Running Incidents — A Step-by-Step Framework

From Chaos to Control — The Four-Step Incident Response Framework Every Team Needs


The Incident Response Framework

When an incident strikes, chaos is the default state. The goal of incident response is to move from chaos to control as quickly as possible.

This framework provides four steps to do exactly that:

  1. Contain the chaos
  2. Stop the bleeding
  3. Communicate like a pro
  4. Close and capture

Step 1: Contain the Chaos

The first step is to establish control over the incident response.

Actions

Action

Description

Why It Matters

Open a channel

Create a dedicated incident channel (Slack, Teams)

Keeps incident communication separate from normal work

Assign roles

Designate Incident Commander, Tech Lead, Comms Lead, Scribe

Ensures someone is accountable for each aspect

Pin the incident doc

Create and pin a shared document

Centralizes all incident information

Declare severity

Determine and declare the incident severity

Sets expectations and resource allocation

Start the timer

Record the start time

Enables accurate MTTD/MTTR measurement

Templates

Incident Channel Template

text

Welcome to the [Service/System] incident response.

 

?? **Current Status**: Investigating

?? **Impact**: [Brief description]

?? **Severity**: P1/P2/P3/P4

 

**Roles**

?? Incident Commander: [Name]

??? Tech Lead: [Name]

?? Comms Lead: [Name]

?? Scribe: [Name]

 

**Incident Doc**: [Link]

 

Please keep all communication in this channel. Use the incident doc for updates.


Step 2: Stop the Bleeding

The second step is to stop the immediate damage.

Actions

Action

Description

Example

Use pre-approved playbooks

Execute known good responses

Roll back a deployment

Roll back changes

Revert recent changes that may have caused the incident

Roll back the last deployment

Flip feature flags

Disable problematic features

Disable the new feature

Fail over

Route traffic to backup systems

Switch to secondary region

Resource scaling

Add capacity to meet demand

Scale up instances

The "First Five Moves"

  1. Get eyes on the issue — Confirm the system is actually broken and to what extent
  2. Identify if any obvious fixes are available — Is this a known issue with a known fix?
  3. Find the blast radius — Are 10 users impacted or 10,000?
  4. Decide on a fix strategy — Is this a rollback or a patch?
  5. Execute the fix — Prefer rollback over new fixes (rollbacks are more predictable)

Step 3: Communicate Like a Pro

The third step is to keep everyone informed.

Actions

Action

Description

Frequency

Internal updates

Update the incident channel

Every 15-30 minutes

External updates

Update the status page

At least every 30 minutes

Exec updates

Brief leadership

At regular intervals

Customer updates

Brief affected customers

As needed

Communication Template

Status Update Template

text

?? **Update #3** | [Time]

 

?? **Status**: Mitigation in progress

 

**What's Happened**:

- [Brief summary of the issue]

- [What's been attempted]

- [What the current plan is]

 

**Impact**:

- [Number of users affected]

- [Business impact]

 

**Next Update**: [Time]

 

**Questions?** : [Name] in the incident channel.

Communication Principles

  1. Be honest: Say what you know and what you don't know
  2. Be timely: Updates every 15-30 minutes
  3. Be clear: Use plain language, not jargon
  4. Be consistent: Same message across all channels
  5. Be accountable: Own the problem and the response

Step 4: Close and Capture

The fourth step is to learn from the incident.

Actions

Action

Description

Timing

Verify service restoration

Confirm the service is fully restored

Immediately

Capture the timeline

Document what happened when

As soon as possible

Schedule the postmortem

Schedule the blameless postmortem

Within 5 business days

Assign action items

Assign owners and dates for improvements

During postmortem

Postmortem Template

Blameless Postmortem

Incident Summary

  • Date: [Date]
  • Start Time: [Time]
  • Resolution Time: [Time]
  • Duration: [Duration]
  • Severity: [P1/P2/P3/P4]
  • Impact: [Description of user impact]

What Happened

  • [Timeline of events]

Why It Happened

  • [Root cause analysis]

What Worked

  • [What went well in the response]

What Could Have Been Better

  • [Areas for improvement]

Action Items

#

Action Item

Owner

Due Date

1

[Action]

[Name]

[Date]

2

[Action]

[Name]

[Date]

Blameless Principle
This postmortem focuses on what happened and how to prevent recurrence, not on who caused it. We recognize that systems fail, not people.


The Incident Response Maturity Model

Level

Description

Characteristics

Level 1: Ad Hoc

No formal incident response

Chaos, heroics, no documentation

Level 2: Defined

Basic incident response

Roles defined, some documentation

Level 3: Managed

Consistent incident response

Regular practice, consistent execution

Level 4: Optimized

Continuous improvement

Learning from every incident, proactive prevention


Conclusion: From Chaos to Control

Incident response doesn't have to be chaotic. With a clear framework, defined roles, and consistent practice, teams can move from chaos to control quickly and effectively.

The framework is simple: Contain, Stop, Communicate, Close. But executing it well requires practice, preparation, and commitment.


Action Items for Your Organization

  • Build runbooks: Create pre-approved playbooks for common incident types
  • Train teams: Practice incident response regularly
  • Conduct postmortems: Learn from every incident
  • Use templates: Provide incident templates for communication, documentation, and postmortems
  • Measure performance: Track MTTD, MTTR, and other incident metrics