Project challenges / verified progress
Beacon: run it like an SRE

The engineering notebook

Run the bad hour

What should the team do during the hour when Beacon is hurting users?

Loading statusStage 9 of 10

  • Workspace not ready
  • Agent not ready
Focus25:00
A small focus ritual

0 focus sessions completed. Every fourth session offers a longer break. Start each phase when you are ready.

Study time never unlocks verified lesson progress.

Loading...

Loading verified progress...

Loading GitHub account...
Phase 4 — Operate and learn

Step 01 of 06

Learn the concept

The first minutes of an incident are where teams lose the plot. Someone debugs, someone posts guesses, someone asks whether it is really an incident. A declared structure is slower for ten seconds and faster for the next hour.

THE FIRST HOURminutesDeclarename commander0-5Stabilizemitigate impact5-20Communicatecadence and scope20-45Handoffnext actions45-60
The timeline is not a script; it is a bias. Declare, assign roles, reduce harm, communicate on a cadence, and preserve the record. Deep diagnosis is valuable after users are safer.
Step 01

The ideas this is made of

Declaration creates permission to coordinate

People declare too late because the word feels dramatic. Treat declaration as a coordination tool, not a confession of failure. It opens the incident channel, assigns roles and starts the timeline. If the alert clears quickly, close it with a note. If it grows, the team is already organized instead of retrofitting process during panic.

The commander owns the process, not the keyboard

An incident commander keeps the goal clear, assigns work, tracks time and asks what decision is needed next. The ops lead changes systems. Communications writes updates. Combining all three roles works for tiny teams only until the first hard incident. Separation prevents the best debugger from also becoming the bottleneck for every status question.

Mitigation beats beautiful diagnosis

If a new release caused Beacon checks to fail, rollback may be correct before anyone understands the exact nil pointer. If an external dependency is slow, reduce concurrency or disable a feature before reading every trace. Root cause matters, but users experience duration. A mitigated incident gives the team oxygen to diagnose accurately.

Severity is impact plus urgency

A bug affecting one internal dashboard is not SEV1 because it is technically fascinating. A failure delaying certificate-expiry alerts for every user may be. Severity should combine user impact, scope, duration and workaround. Decide it explicitly and revise it when facts change. The severity drives communication cadence and leadership attention.

A terse incident update
14:05 UTC — SEV2 declared for Beacon check delays.
Impact: scheduled endpoint checks are running 8-12 minutes late for most users.
Mitigation: canary release rolled back; queue drain in progress.
Next update: 14:20 UTC, or earlier if user impact changes.

The update states impact, action and next communication time. It does not narrate every hypothesis. That restraint keeps responders working and stakeholders informed.

Incident roles

RoleOwnsAvoids

Commander

Coordination and decisions

Deep debugging

Ops lead

System changes

Status writing

Comms

Updates and timeline

Changing prod

Scribe

Facts and timestamps

Interpretation wars

What these are called on the job

  • SEV — Severity level, usually based on impact, scope and urgency.

  • Mitigation — An action that reduces user harm, even before root cause is known.

  • Timeline — Timestamped record of alerts, decisions, actions and observations.