Step 01 of 06
Learn the concept
Every noisy alert trains humans to distrust the next one. That is not inconvenience; it is a reliability risk. Beacon should page when users are burning budget, not when a graph makes a funny shape.
The ideas this is made of
Cause alerts are guesses with a siren
CPU high, pod restarted, queue length up and disk I/O slow may be useful diagnostic signals. They are poor pages unless they reliably imply user harm. Cause alerts fire during harmless maintenance, expected bursts and self-healing events. Symptom alerts start with the user's pain: error budget burning, requests failing, latency above the SLO threshold.
Burn rate says how fast the month is being spent
A burn rate of 1 means the service is consuming budget exactly at the allowed long-term pace. A burn rate of 14.4 against a 99.9% SLO means a full 30-day budget would be exhausted in about 50 hours. That is worth attention. The alert threshold is not magic; it is arithmetic tied to the budget.
Two windows separate disasters from noise
A one-hour fast-burn window catches severe outages quickly, but it can be sensitive to short spikes. A six-hour slow-burn window catches persistent smaller problems, but it would react too slowly to a total outage. Multi-window rules require both a short and a longer view for the same class of burn, reducing false positives without hiding real harm.
Alert fatigue is a system failure
When a team receives pages that require no action, it adapts. People mute phones, delay acknowledgement or treat the alert as probably false. The next real outage pays for that noise. Alert quality is not politeness; it is part of the reliability design. Every page should demand a human decision now.
Target: 99.9% over 30 days
Budget: 0.1% of valid events
Fast page: spend 2% of budget in 1 hour
Allowed monthly spend per hour: 1 / (30 * 24) = 0.00139 budgets/hour
Fast burn threshold: 0.02 / 1 = 14.4 times normal
Slow page: spend 5% in 6 hours
Slow burn threshold: 0.05 / 6 / 0.00139 = 6.0 times normalThese are the classic SRE thresholds because they tie urgency to budget loss. The numbers can change, but the reasoning should not.
Cause alerts versus symptom alerts
| Alert | What it means | Page? |
|---|---|---|
CPU > 90% | A resource is busy | Usually no |
Pod restarted | A process restarted | Usually no |
5% budget in 6h | Users are steadily hurt | Yes |
2% budget in 1h | Users are acutely hurt | Yes |
What these are called on the job
Page — An alert that interrupts a human immediately.
Ticket alert — An alert that needs action but not immediate interruption.
forduration — How long a Prometheus alert condition must remain true before firing.
