Project challenges / verified progress
Beacon: run it like an SRE

The engineering notebook

Close the loop

How does Beacon turn one bad night into a more reliable system?

Loading statusStage 10 of 10

  • Workspace not ready
  • Agent not ready
Focus25:00
A small focus ritual

0 focus sessions completed. Every fourth session offers a longer break. Start each phase when you are ready.

Study time never unlocks verified lesson progress.

Loading...

Loading verified progress...

Loading GitHub account...
Phase 4 — Operate and learn

Step 01 of 06

Learn the concept

The final loop is the quiet one. Beacon was built to warn people before their endpoints fail silently; now Beacon has its own alert, its own dashboard and its own runbook. A system is not operated because it runs. It is operated because it learns.

BEACON WATCHES BEACON01Instrumentlatency and errorsmetrics02EvaluateSLO and budgetrules03Notifyroute with contextalert04Respondrunbook firstincident05Learnpostmortem actionsclose loop
The course ends where Beacon began: a measured fact becomes a useful warning. The difference is direction. Beacon no longer only watches other endpoints; its own operation is visible, budgeted, alerted, routed and improved.
Step 01

The ideas this is made of

A runbook is executable prose

A good runbook names symptoms, confirms scope, lists safe mitigations, gives rollback commands or links, identifies owners and states when to escalate. It assumes the reader is tired and unfamiliar with the code. It should not say investigate logs without naming which logs, where they live and what patterns matter. Ambiguity is toil with nicer typography.

Blameless does not mean consequence-free

Blameless postmortems reject Alice clicked the wrong button as a root cause. They ask why the button was dangerous, why review did not catch it, why rollback was slow and why the alert was late. People remain accountable for actions; the organization becomes accountable for designing systems where ordinary human behaviour is less likely to cause harm.

Contributing factors beat a single root cause

Most incidents are braids, not arrows. A risky deploy, missing dashboard, vague runbook and overloaded on-call can all contribute. The contributing-factors model keeps the team from stopping at the first satisfying explanation. It also produces better fixes: remove several weak conditions rather than pretending one perfect patch will prevent a repeat.

Action items need a test of completion

Improve alerting is not an action item. Add a fast-burn alert for check success, owner Priya, due Friday, verified by promtool and a synthetic failure is. Owners and dates create accountability; verification prevents performative paperwork. The postmortem is not complete when the document is written. It is complete when the learning is installed.

A postmortem action with teeth
- Action: Add a recording rule for six-hour Beacon SLO burn and display it on the overview dashboard.
  Owner: Samir
  Due: 2026-10-12
  Verification: promtool check rules passes, dashboard panel shows non-empty data in local compose.
  Linked factor: slow-burn degradation was visible only after manual PromQL during the incident.

The action is specific, assigned, dated and testable. It names the contributing factor it removes, so future reviewers can tell whether the fix matches the lesson.

Weak versus usable learning

ArtifactWeakUsable

Runbook

Check the logs

Open this dashboard, run this query

Cause

Human error

Risky UI plus missing guardrail

Action

Improve tests

Owner, date, verification

Closure

Doc published

Fix shipped and checked

What these are called on the job

  • Runbook — A procedure for responding to a known operational condition.

  • Postmortem — A written analysis of an incident's impact, timeline, contributing factors and follow-up actions.

  • Blameless — Focused on system conditions and learning rather than individual punishment.

  • Corrective action — A specific change that reduces recurrence or improves detection and response.