A typical P1 incident breaks down like this: 12 minutes assembling the team and gathering context, 20 minutes troubleshooting, 4 minutes on mitigation, 12 minutes on cleanup. The fix - the part engineers actually train for - is the smallest slice. The two brackets around it, the opening scramble and the closing cleanup, are where incidents quietly bleed time. Both are coordination problems, not engineering problems. And both happen in Slack.
This is a playbook for the on-call handoff specifically: the moment shift changes while something is still live, or a new engineer picks up a channel where the previous responder left implicit context that never got written down.
Why on-call handoffs fail and where the time goes
Coordination consumes roughly 70% of total resolution time, while the actual repair takes a fraction of it. That number sounds extreme until you see the breakdown. The typical workflow spans five separate tools: PagerDuty for alerting, Datadog for metrics, Slack for comms, Google Docs for notes, Jira for tickets - five tools, twelve minutes of logistics before troubleshooting starts.
The handoff problem is a second instance of the same pattern. When shift changes mid-incident, the incoming engineer faces the same cold start that the first responder faced at the alert, except now the relevant context is scattered across a channel thread instead of being assembled in one place. Escalating to the right owner takes time, and each handoff resets part of the investigation.
A less-discussed figure: the target for well-run teams is fewer than 5% of resolved incidents reopened within 4 hours of handoff. Most teams don't track this number. The ones that do find it's a direct indicator of handoff quality, not engineering quality. The incident wasn't misclosed; the incoming engineer just didn't have the context to hold the resolution.
The three posts every incident channel needs
An incident channel in Slack is not a chat thread. It's a shared operating surface. Keep the main channel for decisions and status; use threads for technical detail. Here are the three structured posts that carry the channel through its lifecycle.
1. The declare post (posted the moment the incident is named)
This goes at the top and stays pinned. It should cover:
- Severity (P1/P2/P3) and the service or feature affected
- First-noticed time and how it was detected (alert, user report, internal)
- Incident Commander name and @mention
- Comms lead (who is updating the status page / external stakeholders)
- Current working hypothesis - even "unknown" is a valid entry
- Link to the runbook for this service
Keep updates concise and structured - impact (who's affected and how), scope (estimated reach), action (what's been tried), ETA (when resolution is expected). That four-field format applies to every subsequent update too.
2. The status update (every 15-30 minutes during the active window)
One message, timestamped, three lines maximum. Example: