Project challenges / verified progress
Beacon: run it like an SRE

The engineering notebook

Open the right dashboard

What dashboard helps during an incident instead of decorating a wall screen?

Loading statusStage 4 of 10

  • Workspace not ready
  • Agent not ready
Focus25:00
A small focus ritual

0 focus sessions completed. Every fourth session offers a longer break. Start each phase when you are ready.

Study time never unlocks verified lesson progress.

Loading...

Loading verified progress...

Loading GitHub account...
Phase 2 — Turn measurements into promises

Step 01 of 06

Learn the concept

A dashboard nobody opens during an incident is office wallpaper. The first useful dashboard answers one question fast: are users hurt, and where should the responder look next? Beacon needs that board before it needs twelve colours of CPU.

ONE SCREEN, TWO METHODSIncident viewfirst page openedREDrate errors durationservicesUSEutil saturation errorsresourcesLinksalerts and runbooksaction
RED tells whether Beacon is serving users well. USE tells whether the machine or cluster is constraining it. The runbook link turns a graph into the next action, which is the point of opening the page at 3am.
Step 01

The ideas this is made of

Dashboards exist to shorten decisions

A panel is useful when someone can name the decision it supports: roll back, scale out, silence a duplicate alert, call the dependency owner. A panel that merely looks technical steals attention during a failure. During an incident, a responder scans for change, scope and impact. Design the first dashboard around those verbs, not around every metric available.

RED starts from the service contract

Rate says how much work Beacon is doing. Errors say how much work failed. Duration says how slow successful or valid work feels. Those three match the user's experience better than CPU does. A service can have 10% CPU and still be broken if every request returns 500. RED keeps the dashboard honest about product behaviour.

USE explains resource pressure without guessing

Utilization asks how busy a resource is. Saturation asks whether work is queued because the resource is too busy. Errors ask whether the resource is failing outright. CPU at 95% may be fine if there is no saturation and latency is stable. A full queue with moderate CPU is a better signal that users are waiting.

Provisioning makes dashboards reviewable

A hand-edited Grafana dashboard is production configuration with no history. Provisioning stores JSON and data-source YAML in the repository, so a dashboard change gets reviewed like code. It also makes local rebuilds boring: start Grafana, load the same dashboard, investigate with the same panels. Boring is praise in operations.

A panel definition with a job
{
  "title": "Beacon error ratio",
  "type": "timeseries",
  "targets": [
    {
      "expr": "sum(rate(beacon_checks_total{result="failure"}[5m])) / sum(rate(beacon_checks_total[5m]))",
      "legendFormat": "failure ratio"
    }
  ],
  "fieldConfig": { "defaults": { "unit": "percentunit", "thresholds": { "mode": "absolute" } } }
}

The panel has a question, a unit and a query tied to Beacon's service behaviour. It is not a random metric dump. A responder can compare it with the SLO panels later.

USE and RED answer different questions

MethodSubjectSignals

RED

A service

Rate, errors, duration

USE

A resource

Utilization, saturation, errors

RED failure

Users hurt

500s or slow requests

USE failure

Capacity hurts

Queues, throttling, I/O errors

What these are called on the job

  • Provisioning — Loading Grafana data sources and dashboards from files at startup.

  • Panel — One dashboard visualization with one operational question.

  • Runbook link — A URL from a panel or alert to the procedure a responder should follow.