Step 01 of 06
Learn the concept
The first minutes of an incident are where teams lose the plot. Someone debugs, someone posts guesses, someone asks whether it is really an incident. A declared structure is slower for ten seconds and faster for the next hour.
The ideas this is made of
Declaration creates permission to coordinate
People declare too late because the word feels dramatic. Treat declaration as a coordination tool, not a confession of failure. It opens the incident channel, assigns roles and starts the timeline. If the alert clears quickly, close it with a note. If it grows, the team is already organized instead of retrofitting process during panic.
The commander owns the process, not the keyboard
An incident commander keeps the goal clear, assigns work, tracks time and asks what decision is needed next. The ops lead changes systems. Communications writes updates. Combining all three roles works for tiny teams only until the first hard incident. Separation prevents the best debugger from also becoming the bottleneck for every status question.
Mitigation beats beautiful diagnosis
If a new release caused Beacon checks to fail, rollback may be correct before anyone understands the exact nil pointer. If an external dependency is slow, reduce concurrency or disable a feature before reading every trace. Root cause matters, but users experience duration. A mitigated incident gives the team oxygen to diagnose accurately.
Severity is impact plus urgency
A bug affecting one internal dashboard is not SEV1 because it is technically fascinating. A failure delaying certificate-expiry alerts for every user may be. Severity should combine user impact, scope, duration and workaround. Decide it explicitly and revise it when facts change. The severity drives communication cadence and leadership attention.
14:05 UTC — SEV2 declared for Beacon check delays.
Impact: scheduled endpoint checks are running 8-12 minutes late for most users.
Mitigation: canary release rolled back; queue drain in progress.
Next update: 14:20 UTC, or earlier if user impact changes.The update states impact, action and next communication time. It does not narrate every hypothesis. That restraint keeps responders working and stakeholders informed.
Incident roles
| Role | Owns | Avoids |
|---|---|---|
Commander | Coordination and decisions | Deep debugging |
Ops lead | System changes | Status writing |
Comms | Updates and timeline | Changing prod |
Scribe | Facts and timestamps | Interpretation wars |
What these are called on the job
SEV — Severity level, usually based on impact, scope and urgency.
Mitigation — An action that reduces user harm, even before root cause is known.
Timeline — Timestamped record of alerts, decisions, actions and observations.
