Project challenges / verified progress
Beacon: run it like an SRE

The engineering notebook

Route the page

How does an alert reach the right place with enough context to act?

Loading statusStage 8 of 10

  • Workspace not ready
  • Agent not ready
Focus25:00
A small focus ritual

0 focus sessions completed. Every fourth session offers a longer break. Start each phase when you are ready.

Study time never unlocks verified lesson progress.

Loading...

Loading verified progress...

Loading GitHub account...
Phase 3 — Page only for real harm

Step 01 of 06

Learn the concept

An alert rule has not helped anyone until it reaches a human who can act. Sending every alert to one channel is not simplicity; it is a fog machine. Alertmanager turns firing rules into routed, grouped, inhibited notifications.

ONE ALERT, MANY DESTINATIONSFiring alertlabels and annotationsPage routeseverity equals pageDiscord nowTicket routeseverity equals ticketworkdaySilencedplanned maintenanceno noise
Routing is label-driven. The same Alertmanager can wake someone for fast SLO burn, create a work item for slow burn, and suppress secondary noise during maintenance. The receiver changes; the alert fact remains the same.
Step 01

The ideas this is made of

Grouping reduces duplicate panic

If ten Beacon replicas all report the same SLO burn, responders need one incident, not ten messages. Alertmanager groups alerts by labels such as alertname and service, waits briefly for related alerts, then sends a combined notification. Good grouping preserves scope while avoiding notification storms. Bad grouping hides separate incidents in one blob.

Inhibition encodes which alerts are secondary

When a high-level Beacon SLO page is firing, a dozen lower-level latency or pod alerts may add no value. Inhibition lets Alertmanager suppress those secondary notifications when labels match. It does not delete alerts; it stops duplicate messages. Use it carefully: inhibiting the only symptom alert with a flaky cause alert is how silence becomes outage camouflage.

Silences are for expected noise, not shame

A silence says: this alert is expected between these times for this label set. It should have a creator, comment and expiry. Permanent silences are abandoned policy. During maintenance, silences keep trust in the pager by making planned work explicit rather than forcing responders to mentally subtract known noise.

Annotations are written for a tired responder

At 3am, expr=(1 - ratio) > 0.0144 is not a message. A useful alert says the impact, when it started, where to look and what runbook to follow. Include dashboard and runbook links. Put the metric name in the details, not as the entire explanation. The receiver should know the first move without reading PromQL.

A humane alert annotation
annotations:
  summary: Beacon check success is burning budget
  impact: Users may receive late or missing endpoint failure and certificate-expiry checks.
  dashboard: http://localhost:3000/d/beacon-overview
  runbook: docs/runbooks/beacon-slo-burn.md
  first_action: Check the RED dashboard, then pause risky deploys if the burn is real.

The metric is not the message. The annotation tells the responder why they were contacted and what to open first.

Alertmanager tools

ToolPurposeDanger

Route

Choose receiver

Labels too vague

Group

Combine related alerts

Hide separate failures

Inhibit

Suppress secondary noise

Mask useful symptoms

Silence

Mute expected alerts

Permanent black hole

What these are called on the job

  • Receiver — A destination such as Discord, email, PagerDuty or a webhook.

  • Route — A label-matching branch in Alertmanager's notification tree.

  • Inhibition — Suppressing one alert notification while another related alert is active.