Project challenges / verified progress
Beacon: run it like an SRE

The engineering notebook

Page on symptoms

When is an alert worth waking a human?

Loading statusStage 7 of 10

  • Workspace not ready
  • Agent not ready
Focus25:00
A small focus ritual

0 focus sessions completed. Every fourth session offers a longer break. Start each phase when you are ready.

Study time never unlocks verified lesson progress.

Loading...

Loading verified progress...

Loading GitHub account...
Phase 3 — Page only for real harm

Step 01 of 06

Learn the concept

Every noisy alert trains humans to distrust the next one. That is not inconvenience; it is a reliability risk. Beacon should page when users are burning budget, not when a graph makes a funny shape.

FAST AND SLOW BURNSbudgetFast burn2 percent in 1hpage nowSlow burn5 percent in 6hsame dayHealthy spendeven over 30dno page
Fast burn catches catastrophic user pain before the month is gone. Slow burn catches a steady leak before it becomes the month's story. Both are symptoms because they measure budget consumption, not guessed root causes.
Step 01

The ideas this is made of

Cause alerts are guesses with a siren

CPU high, pod restarted, queue length up and disk I/O slow may be useful diagnostic signals. They are poor pages unless they reliably imply user harm. Cause alerts fire during harmless maintenance, expected bursts and self-healing events. Symptom alerts start with the user's pain: error budget burning, requests failing, latency above the SLO threshold.

Burn rate says how fast the month is being spent

A burn rate of 1 means the service is consuming budget exactly at the allowed long-term pace. A burn rate of 14.4 against a 99.9% SLO means a full 30-day budget would be exhausted in about 50 hours. That is worth attention. The alert threshold is not magic; it is arithmetic tied to the budget.

Two windows separate disasters from noise

A one-hour fast-burn window catches severe outages quickly, but it can be sensitive to short spikes. A six-hour slow-burn window catches persistent smaller problems, but it would react too slowly to a total outage. Multi-window rules require both a short and a longer view for the same class of burn, reducing false positives without hiding real harm.

Alert fatigue is a system failure

When a team receives pages that require no action, it adapts. People mute phones, delay acknowledgement or treat the alert as probably false. The next real outage pays for that noise. Alert quality is not politeness; it is part of the reliability design. Every page should demand a human decision now.

Burn-rate threshold from an SLO
Target: 99.9% over 30 days
Budget: 0.1% of valid events

Fast page: spend 2% of budget in 1 hour
Allowed monthly spend per hour: 1 / (30 * 24) = 0.00139 budgets/hour
Fast burn threshold: 0.02 / 1 = 14.4 times normal

Slow page: spend 5% in 6 hours
Slow burn threshold: 0.05 / 6 / 0.00139 = 6.0 times normal

These are the classic SRE thresholds because they tie urgency to budget loss. The numbers can change, but the reasoning should not.

Cause alerts versus symptom alerts

AlertWhat it meansPage?

CPU > 90%

A resource is busy

Usually no

Pod restarted

A process restarted

Usually no

5% budget in 6h

Users are steadily hurt

Yes

2% budget in 1h

Users are acutely hurt

Yes

What these are called on the job

  • Page — An alert that interrupts a human immediately.

  • Ticket alert — An alert that needs action but not immediate interruption.

  • for duration — How long a Prometheus alert condition must remain true before firing.