The Single Most Counterintuitive Rule of Modern Incident Response
The Greatest Asset Becomes the Greatest Liability
You have a P1 incident. The system is down. Customers are screaming. Revenue is being lost.
Your instinct is obvious: get your best engineer on the problem. They know the system inside out. They've saved the day before. They can fix this.
But your best engineer is also your worst enemy right now.
The Incident Commander model contains the single most counterintuitive rule of modern incident response: the Incident Commander does not touch the keyboard.
Why the Best Engineer Shouldn't Be in the Driver's Seat
When your best engineer starts debugging:
- They lose oversight: They become focused on a single issue, missing the bigger picture
- They stop coordinating: The team loses its leader
- They stop communicating: Stakeholders don't know what's happening
- They become a bottleneck: Everyone waits for their findings
- They burn out: They're doing two jobs—technical leadership and technical execution
The moment the Incident Commander starts debugging code, they lose oversight, and that's when things cascade.
The Four Key Incident Response Roles
Role 1: Incident Commander
|
Dimension |
Description |
|
Responsibility |
Owns the response, makes decisions, keeps the team focused |
|
Action |
Asks sharp questions, sets priorities, delegates tasks, keeps the timeline moving |
|
Prohibited |
Does NOT touch the keyboard |
|
Skills |
Decision-making, communication, delegation, situational awareness |
Role 2: Technical Lead / Operations Lead
|
Dimension |
Description |
|
Responsibility |
Investigates root cause, proposes and executes fixes |
|
Action |
Leads technical investigation, implements fixes |
|
Skills |
Deep technical expertise, problem-solving |
Role 3: Communications Lead
|
Dimension |
Description |
|
Responsibility |
Updates internal stakeholders and status page |
|
Action |
Manages internal and external communications |
|
Skills |
Communication, clarity, calm under pressure |
Role 4: Scribe
|
Dimension |
Description |
|
Responsibility |
Captures timeline, key decisions, and actions taken |
|
Action |
Documents everything in real-time |
|
Skills |
Attention to detail, documentation |
The Five-Stage Incident Response Lifecycle
Many SRE teams have evolved beyond NIST's four-phase framework to a five-stage model:
- Prepare: Build systems, runbooks, and teams before incidents happen
- Detect: Identify that an incident is occurring
- Respond: Mobilize and coordinate the response
- Recover: Restore service and verify resolution
- Learn: Conduct blameless postmortems and improve
The Incident Commander is critical in the Respond and Recover stages.
The Incident Commander's Responsibilities
During an incident, the Incident Commander:
Starts the Response
- Opens an incident channel
- Assigns roles
- Sets the initial priorities
Coordinates the Response
- Delegates tasks
- Manages resources
- Keeps the team focused
Communicates
- Updates stakeholders
- Manages external communications
- Keeps the status page current
Makes Decisions
- Makes decisions when the team is uncertain
- Sets priorities
- Escalates when necessary
Closes the Incident
- Verifies service is restored
- Captures the timeline
- Schedules the postmortem
The "First Five Moves" for Incident Mitigation
When an incident strikes, the Incident Commander should execute the "first five moves":
- Open an incident channel
- Assign roles (IC, Ops Lead, Comms Lead, Scribe)
- Pin the incident doc (where all information will be captured)
- Declare the severity (P1, P2, etc.)
- Start communicating (stakeholders, status page)
Building Incident Command Capabilities
1. Define Roles Clearly
Document exactly what each role does during an incident. Include:
- Responsibilities
- Actions
- Escalation paths
2. Train Your Team
Conduct regular training on incident command:
- Role-playing
- Tabletop exercises
- Live incident observation
3. Use Runbooks
Pre-approved playbooks for common incident types:
- Roll back changes
- Flip feature flags
- Fail over to backup systems
4. Conduct Postmortems
Blameless postmortems within 5 business days:
- What happened
- Why it happened
- What we'll do differently
The Culture Shift
The Incident Commander model requires a culture shift:
From Hero Culture to System Culture
|
Dimension |
Hero Culture |
System Culture |
|
Who's responsible |
Individuals |
Systems |
|
What's celebrated |
Heroic fixes |
Effective processes |
|
Response |
Individual heroics |
Coordinated team response |
|
Postmortem |
Blame |
Learning |
From Blame to Learning
Blameless postmortems are the foundation of incident mastery:
- Focus on what failed in the system, not who caused it
- Turn findings into action items
- Schedule postmortems within 5 business days
Conclusion: The Best Engineer Leads, Doesn't Execute
The Incident Commander model is counterintuitive: the best engineer shouldn't touch the keyboard. But it's also effective: coordinated response beats individual heroics every time.
Your best engineer's greatest value during an incident isn't their keyboard skills. It's their judgment, their decision-making, and their ability to lead.
Action Items for Your Organization
- Define incident roles: Document Incident Commander, Technical Lead, Comms Lead, Scribe
- Train your team: Conduct regular incident command training
- Build runbooks: Create pre-approved playbooks for common incident types
- Conduct tabletop exercises: Practice incident response
- Establish blameless postmortems: Focus on learning, not blame
- Document everything: Capture incident timelines, decisions, and actions