The scribe is the first role to collapse under pressure. On a small team handling a SEV-1 at 2am, the person nominally assigned to log decisions quietly shifts to debugging, and by the time the incident is over, teams reconstruct timelines from fragmented Slack histories - and often get them wrong. That reconstruction is where postmortems go thin and runbooks stay stale.
This playbook covers the full arc: from the moment someone types /incident to the moment the postmortem actions land in a ticket. It is written for teams already living in Slack, handling P1s and P2s at least a few times a month.
The six stages, with target times
A well-run Slack-first incident moves through six stages with explicit time targets for P1s: Detect (signal reaches a human in under 2 minutes), Declare (incident exists with priority and owner in under 5 minutes), Assemble (right people on the bridge in under 10 minutes), Communicate (first stakeholder update sent in under 15 minutes), Resolve (service restored and verified per SLA), and Review (PIR drafted and actions assigned within 5 business days).
Most teams have the first three stages reasonably wired - alerting fires, PagerDuty pages, a channel appears. The failures cluster in stages 4 through 6: the stakeholder update that goes out 40 minutes late, the scribe who stopped logging at minute 20, the postmortem that gets drafted from memory a week later.
Elite teams achieve MTTR under 60 minutes (DORA Report); low performers average over 24 hours. The spread is not primarily about engineering skill. It is about the quality of coordination inside the incident channel while the fix is being worked.
The three roles you must assign in the first five minutes
The incident channel is the room. Assign three roles immediately: the Incident Commander (IC), who owns decisions and timelines and does not debug; the Technical Lead, who coordinates engineers working the fix; and the Comms Lead, who owns stakeholder updates (often the IC on smaller teams).
The fourth role - scribe - is rarely listed in the same breath, but it deserves to be. The scribe maintains a timestamped record of every action taken, decision made, and update sent throughout the incident, and this record becomes the foundation of the blameless postmortem.
In teams of fewer than six engineers, the scribe and comms lead roles are often combined - workable for shorter incidents, but a risk for anything running longer than two hours. A teammate like Beagle is better suited to the scribe role than any human on a small team under load: it watches the channel, timestamps decisions as they land, and does not defect to debugging when things get tense.
The update cadence that kills the most inbound pings
The single most effective communication practice - confirmed across incident.io, PagerDuty, and Rootly research - is always stating when the next update will arrive. Close every update with "Next update by [time]." This one sentence eliminates the majority of inbound stakeholder pings during active incidents, because it removes uncertainty about whether anyone is handling communication at all.
What the update should contain varies by audience. Internal and external messages are different documents.
| Dimension | Internal (engineers + leadership) | External (customers + partners) |
|---|---|---|
| Channel | Incident Slack channel | Status page + email |
| Tone | Technical, factual | Plain language, impact-focused |
| Content | Root cause hypothesis, next steps | What's affected and when to expect resolution |
| Who sends | IC or Comms Lead | Comms Lead, PR review for SEV1 |
| Cadence (SEV1) | Every 20-30 min | Every 30-60 min |
The two biggest communication failures during incidents are silence and speculation. Silence generates inbound pings. Speculation - leaking an internal root-cause guess to an external audience - generates support calls. Pick one place that holds the official version of events (almost always a status page). Every other channel points back to that source. If the status page says "Investigating," support cannot tell a customer "resolved in 10 minutes," and the Slack channel cannot leak an internal guess about root cause to an external audience.
Channel hygiene that keeps the main thread readable
Dedicated channels create focus. A single channel per incident means everyone involved sees the same information - no cross-talk, no "did you see my message in #engineering?" The channel IS the incident.
Inside the channel, discipline on thread use is what separates a readable incident log from a wall of noise.
Main channel: declarations, role assignments, status updates, resolution notice. One post per decision.
Threads: debugging details, log dumps, side investigations. Threads keep debugging details and log dumps out of the main channel, and using reactions instead of status messages reduces noise.
Channel naming: use consistent prefixes -
incident-,alert-,sev1-followed by date and a brief description (e.g.,
#inc-2026-1007-payments). The name is how you find this channel in six months when a similar pattern fires.For P3 and P4: a dedicated channel per incident is warranted for P1 and P2; a thread in a shared channel is usually enough for P3 and P4.
Closing the loop: resolution and the PIR
Resolution is not the end of the incident - it is the end of the acute phase. Two things need to happen before the channel is archived.
First, post an explicit resolution notice: what was fixed, when service was verified restored, and a link to the postmortem (even if it is just a stub). Going silent during a long incident is a mistake; a short "no change, still working on it" update on schedule beats no update at all. The same principle applies at resolution: a "resolved" message with no follow-up leaves customers without confidence the issue won't recur.
Second, draft the postmortem while the channel is still warm. Start the postmortem during incident resolution. Capture decisions, assumptions, and logs while they're still fresh. Do not wait until after recovery. The scribe log - if someone kept it - makes this a 20-minute edit rather than a 90-minute reconstruction. If a teammate like Beagle was scribing, the draft is already structured: decisions with timestamps, a causal chain, open action items.
Action item completion rate - the percentage of postmortem actions completed by deadline - is the metric that shows whether learning turns into execution. Most teams track it loosely if at all. A structured log makes it measurable.
Incident channel playbook: common questions
What roles are required in a Slack incident channel?
Three at minimum: Incident Commander (owns decisions, does not debug), Technical Lead (coordinates the fix), and Comms Lead (owns stakeholder updates). A scribe - responsible for timestamped logging - is the fourth role, and the first to get dropped. On teams under six people, IC and Comms Lead are often the same person; scribe should still be assigned separately.
How often should you post updates in an incident channel?
For SEV1/P1 incidents, every 20-30 minutes internally and every 30-60 minutes to external stakeholders. Every update should close with an explicit "next update by [time]" line. This single habit removes most inbound status pings, because stakeholders know exactly when to expect news.
Should every incident get its own Slack channel?
For P1 and P2, yes - a dedicated channel keeps communication clean and produces a searchable log. For P3 and P4, a thread in a shared team channel is usually enough. Naming convention matters: use a consistent prefix (e.g., #inc-) plus date and a brief description so the channel is findable months later.
What goes in the main channel vs. a thread?
Main channel: declarations, role assignments, status updates, resolution notice - one post per significant decision. Threads: debugging detail, log snippets, and side investigations. Keeping the main channel clean means anyone joining mid-incident can read the last 10 posts and understand exactly where things stand.
What makes a postmortem actually usable?
A timestamped scribe log from inside the incident channel. Without it, postmortems get reconstructed from memory and miss the decision points that mattered. The log should capture: when each role was assigned, what hypotheses were tested and discarded, when the fix was applied, and when resolution was verified. That structure maps directly to the postmortem template fields for timeline, root cause, and contributing factors.