The engineering notebook
Beacon: run it like an SRE
You finish able to instrument a Go service with Prometheus metrics, scrape and query it, build useful dashboards, define SLIs and SLOs, alert on burn rate, route notifications, run incidents and write postmortems that change the system.
Your learning trail
Verified project-agent submissions only. Reading or clicking cannot unlock progress.
System design
What you are building
Beacon becomes an operated service by exposing bounded Prometheus metrics, scraping them locally, turning them into PromQL answers, displaying the answers on incident-first dashboards, recording SLO facts, paging on budget burn, routing alerts with context and documenting the response. The system that once only watched external endpoints now watches itself.
Stages
10 stages, in order
Expand any stage to read what it teaches. A stage opens for work once the stage before it passes a verified submission.
Phase 1 — Make Beacon visible
Learn the observability model, instrument Beacon correctly and scrape the resulting evidence with Prometheus.
See the running system75 minutes (locked)
How do you know what a running system is doing before it wakes you up?
You will be able to create deploy/observability/ as the home for local Prometheus, Grafana and Alertmanager configuration.
Fork the project repository to start working through the stages.
Instrument the service90 minutes (locked)
What should Beacon expose so Prometheus can measure Beacon itself?
You will be able to add a metrics package that registers Beacon HTTP request, check result and check duration metrics.
Fork the project repository to start working through the stages.
Ask the data questions2 hours (locked)
How does Prometheus collect Beacon metrics and turn samples into answers?
You will be able to create a local Prometheus configuration under deploy/observability/.
Fork the project repository to start working through the stages.
Phase 2 — Turn measurements into promises
Build dashboards, define user-centred SLIs and turn SLO targets into recorded error-budget math.
Open the right dashboard90 minutes (locked)
What dashboard helps during an incident instead of decorating a wall screen?
You will be able to create Grafana provisioning directories under deploy/observability/grafana/.
Fork the project repository to start working through the stages.
Measure what users feel90 minutes (locked)
Which measurement tells whether Beacon is actually working for its users?
You will be able to write SLI specifications for Beacon check success and Beacon check latency.
Fork the project repository to start working through the stages.
Spend the error budget2 hours (locked)
How does one reliability number become an engineering policy?
You will be able to choose a 30-day rolling SLO target for Beacon check success, such as 99.9%.
Fork the project repository to start working through the stages.
Phase 3 — Page only for real harm
Alert on SLO burn, route notifications with context and keep humans from becoming the weakest reliability link.
Page on symptoms2 hours (locked)
When is an alert worth waking a human?
You will be able to add Prometheus alert rules for fast and slow SLO burn on Beacon check success.
Fork the project repository to start working through the stages.
Route the page90 minutes (locked)
How does an alert reach the right place with enough context to act?
You will be able to create an Alertmanager configuration under deploy/observability/alertmanager/.
Fork the project repository to start working through the stages.
Phase 4 — Operate and learn
Run the incident hour deliberately and close the Beacon loop with usable runbooks and blameless postmortems.
Run the bad hour90 minutes (locked)
What should the team do during the hour when Beacon is hurting users?
You will be able to create an incident runbook for Beacon SLO burn under docs/runbooks/.
Fork the project repository to start working through the stages.
Close the loop2 hours (locked)
How does Beacon turn one bad night into a more reliable system?
You will be able to finish the Beacon SLO-burn runbook with symptoms, confirmation, mitigation, escalation and rollback sections.
Fork the project repository to start working through the stages.
About this path
Beacon exists to warn other people before endpoints fail silently, yet its own health is still mostly implied by deployment success. That is not operations. This course closes the loop: Beacon measures its own latency, errors, SLO burn and incident response with the same seriousness it applies to certificate expiry for others.
What you will learn: Prometheus instrumentation in Go, PromQL and histogram quantiles, SLI and SLO definition, Error budget burn-rate alerting, Incident response and postmortems. Build the project through cumulative challenges with beginner explanations and local verification.
- Level
- Complete beginner to independently building and operating the project
- Format
- 10 cumulative stages. Every stage teaches the concept in full before any code, then gives you the thing to build and the run that proves it works
- Before you start
- You need Docker, kind, kubectl and Go 1.22 or newer. Prometheus, Grafana and Alertmanager all run locally in containers. The notification stage uses a free Discord webhook. There is zero cloud cost and no paid account requirement.
