The Incident Commander Model — Why Your Best Engineer Shouldn't Touch the Keyboard

The Single Most Counterintuitive Rule of Modern Incident Response


The Greatest Asset Becomes the Greatest Liability

You have a P1 incident. The system is down. Customers are screaming. Revenue is being lost.

Your instinct is obvious: get your best engineer on the problem. They know the system inside out. They've saved the day before. They can fix this.

But your best engineer is also your worst enemy right now.

The Incident Commander model contains the single most counterintuitive rule of modern incident response: the Incident Commander does not touch the keyboard.


Why the Best Engineer Shouldn't Be in the Driver's Seat

When your best engineer starts debugging:

  1. They lose oversight: They become focused on a single issue, missing the bigger picture
  2. They stop coordinating: The team loses its leader
  3. They stop communicating: Stakeholders don't know what's happening
  4. They become a bottleneck: Everyone waits for their findings
  5. They burn out: They're doing two jobs—technical leadership and technical execution

The moment the Incident Commander starts debugging code, they lose oversight, and that's when things cascade.


The Four Key Incident Response Roles

Role 1: Incident Commander

Dimension

Description

Responsibility

Owns the response, makes decisions, keeps the team focused

Action

Asks sharp questions, sets priorities, delegates tasks, keeps the timeline moving

Prohibited

Does NOT touch the keyboard

Skills

Decision-making, communication, delegation, situational awareness

Role 2: Technical Lead / Operations Lead

Dimension

Description

Responsibility

Investigates root cause, proposes and executes fixes

Action

Leads technical investigation, implements fixes

Skills

Deep technical expertise, problem-solving

Role 3: Communications Lead

Dimension

Description

Responsibility

Updates internal stakeholders and status page

Action

Manages internal and external communications

Skills

Communication, clarity, calm under pressure

Role 4: Scribe

Dimension

Description

Responsibility

Captures timeline, key decisions, and actions taken

Action

Documents everything in real-time

Skills

Attention to detail, documentation


The Five-Stage Incident Response Lifecycle

Many SRE teams have evolved beyond NIST's four-phase framework to a five-stage model:

  1. Prepare: Build systems, runbooks, and teams before incidents happen
  2. Detect: Identify that an incident is occurring
  3. Respond: Mobilize and coordinate the response
  4. Recover: Restore service and verify resolution
  5. Learn: Conduct blameless postmortems and improve

The Incident Commander is critical in the Respond and Recover stages.


The Incident Commander's Responsibilities

During an incident, the Incident Commander:

Starts the Response

  • Opens an incident channel
  • Assigns roles
  • Sets the initial priorities

Coordinates the Response

  • Delegates tasks
  • Manages resources
  • Keeps the team focused

Communicates

  • Updates stakeholders
  • Manages external communications
  • Keeps the status page current

Makes Decisions

  • Makes decisions when the team is uncertain
  • Sets priorities
  • Escalates when necessary

Closes the Incident

  • Verifies service is restored
  • Captures the timeline
  • Schedules the postmortem

The "First Five Moves" for Incident Mitigation

When an incident strikes, the Incident Commander should execute the "first five moves":

  1. Open an incident channel
  2. Assign roles (IC, Ops Lead, Comms Lead, Scribe)
  3. Pin the incident doc (where all information will be captured)
  4. Declare the severity (P1, P2, etc.)
  5. Start communicating (stakeholders, status page)

Building Incident Command Capabilities

1. Define Roles Clearly

Document exactly what each role does during an incident. Include:

  • Responsibilities
  • Actions
  • Escalation paths

2. Train Your Team

Conduct regular training on incident command:

  • Role-playing
  • Tabletop exercises
  • Live incident observation

3. Use Runbooks

Pre-approved playbooks for common incident types:

  • Roll back changes
  • Flip feature flags
  • Fail over to backup systems

4. Conduct Postmortems

Blameless postmortems within 5 business days:

  • What happened
  • Why it happened
  • What we'll do differently

The Culture Shift

The Incident Commander model requires a culture shift:

From Hero Culture to System Culture

Dimension

Hero Culture

System Culture

Who's responsible

Individuals

Systems

What's celebrated

Heroic fixes

Effective processes

Response

Individual heroics

Coordinated team response

Postmortem

Blame

Learning

From Blame to Learning

Blameless postmortems are the foundation of incident mastery:

  • Focus on what failed in the system, not who caused it
  • Turn findings into action items
  • Schedule postmortems within 5 business days

Conclusion: The Best Engineer Leads, Doesn't Execute

The Incident Commander model is counterintuitive: the best engineer shouldn't touch the keyboard. But it's also effective: coordinated response beats individual heroics every time.

Your best engineer's greatest value during an incident isn't their keyboard skills. It's their judgment, their decision-making, and their ability to lead.


Action Items for Your Organization

  • Define incident roles: Document Incident Commander, Technical Lead, Comms Lead, Scribe
  • Train your team: Conduct regular incident command training
  • Build runbooks: Create pre-approved playbooks for common incident types
  • Conduct tabletop exercises: Practice incident response
  • Establish blameless postmortems: Focus on learning, not blame
  • Document everything: Capture incident timelines, decisions, and actions