Step 01 of 06
Learn the concept
A dashboard nobody opens during an incident is office wallpaper. The first useful dashboard answers one question fast: are users hurt, and where should the responder look next? Beacon needs that board before it needs twelve colours of CPU.
The ideas this is made of
Dashboards exist to shorten decisions
A panel is useful when someone can name the decision it supports: roll back, scale out, silence a duplicate alert, call the dependency owner. A panel that merely looks technical steals attention during a failure. During an incident, a responder scans for change, scope and impact. Design the first dashboard around those verbs, not around every metric available.
RED starts from the service contract
Rate says how much work Beacon is doing. Errors say how much work failed. Duration says how slow successful or valid work feels. Those three match the user's experience better than CPU does. A service can have 10% CPU and still be broken if every request returns 500. RED keeps the dashboard honest about product behaviour.
USE explains resource pressure without guessing
Utilization asks how busy a resource is. Saturation asks whether work is queued because the resource is too busy. Errors ask whether the resource is failing outright. CPU at 95% may be fine if there is no saturation and latency is stable. A full queue with moderate CPU is a better signal that users are waiting.
Provisioning makes dashboards reviewable
A hand-edited Grafana dashboard is production configuration with no history. Provisioning stores JSON and data-source YAML in the repository, so a dashboard change gets reviewed like code. It also makes local rebuilds boring: start Grafana, load the same dashboard, investigate with the same panels. Boring is praise in operations.
{
"title": "Beacon error ratio",
"type": "timeseries",
"targets": [
{
"expr": "sum(rate(beacon_checks_total{result="failure"}[5m])) / sum(rate(beacon_checks_total[5m]))",
"legendFormat": "failure ratio"
}
],
"fieldConfig": { "defaults": { "unit": "percentunit", "thresholds": { "mode": "absolute" } } }
}The panel has a question, a unit and a query tied to Beacon's service behaviour. It is not a random metric dump. A responder can compare it with the SLO panels later.
USE and RED answer different questions
| Method | Subject | Signals |
|---|---|---|
RED | A service | Rate, errors, duration |
USE | A resource | Utilization, saturation, errors |
RED failure | Users hurt | 500s or slow requests |
USE failure | Capacity hurts | Queues, throttling, I/O errors |
What these are called on the job
Provisioning — Loading Grafana data sources and dashboards from files at startup.
Panel — One dashboard visualization with one operational question.
Runbook link — A URL from a panel or alert to the procedure a responder should follow.
