Project challenges / verified progress
Engineering project paths

The engineering notebook

Beacon: run it like an SRE

You finish able to instrument a Go service with Prometheus metrics, scrape and query it, build useful dashboards, define SLIs and SLOs, alert on burn rate, route notifications, run incidents and write postmortems that change the system.

Your learning trail

0 / 10 complete

Verified project-agent submissions only. Reading or clicking cannot unlock progress.

System design

What you are building

Beacon becomes an operated service by exposing bounded Prometheus metrics, scraping them locally, turning them into PromQL answers, displaying the answers on incident-first dashboards, recording SLO facts, paging on budget burn, routing alerts with context and documenting the response. The system that once only watched external endpoints now watches itself.

BEACON / WATCHING THE WATCHERBeacon itselfnow the thing measuredTHE OBSERVABILITY YOU OWN/metricscounters and histogramsPrometheusscrape, store, querySLI and SLOgood events over validBurn-rate alertsymptoms, not causesWHERE IT ENDS UPGrafanaRED and USE, as codeA human, on callwoken only when it countsevery alert carries a runbook, and every incident improves oneTHE MONITOR IS NOW MONITORED, AND THE LOOP IS CLOSED

Stages

10 stages, in order

Expand any stage to read what it teaches. A stage opens for work once the stage before it passes a verified submission.

Phase 1Stages 1–30 of 3 stages verified

Phase 1 — Make Beacon visible

Learn the observability model, instrument Beacon correctly and scrape the resulting evidence with Prometheus.

  • See the running system75 minutes (locked)

    How do you know what a running system is doing before it wakes you up?

    You will be able to create deploy/observability/ as the home for local Prometheus, Grafana and Alertmanager configuration.

    Fork the project repository to start working through the stages.

  • Instrument the service90 minutes (locked)

    What should Beacon expose so Prometheus can measure Beacon itself?

    You will be able to add a metrics package that registers Beacon HTTP request, check result and check duration metrics.

    Fork the project repository to start working through the stages.

  • Ask the data questions2 hours (locked)

    How does Prometheus collect Beacon metrics and turn samples into answers?

    You will be able to create a local Prometheus configuration under deploy/observability/.

    Fork the project repository to start working through the stages.

Phase 2Stages 4–60 of 3 stages verified

Phase 2 — Turn measurements into promises

Build dashboards, define user-centred SLIs and turn SLO targets into recorded error-budget math.

  • Open the right dashboard90 minutes (locked)

    What dashboard helps during an incident instead of decorating a wall screen?

    You will be able to create Grafana provisioning directories under deploy/observability/grafana/.

    Fork the project repository to start working through the stages.

  • Measure what users feel90 minutes (locked)

    Which measurement tells whether Beacon is actually working for its users?

    You will be able to write SLI specifications for Beacon check success and Beacon check latency.

    Fork the project repository to start working through the stages.

  • Spend the error budget2 hours (locked)

    How does one reliability number become an engineering policy?

    You will be able to choose a 30-day rolling SLO target for Beacon check success, such as 99.9%.

    Fork the project repository to start working through the stages.

Phase 3Stages 7–80 of 2 stages verified

Phase 3 — Page only for real harm

Alert on SLO burn, route notifications with context and keep humans from becoming the weakest reliability link.

  • Page on symptoms2 hours (locked)

    When is an alert worth waking a human?

    You will be able to add Prometheus alert rules for fast and slow SLO burn on Beacon check success.

    Fork the project repository to start working through the stages.

  • Route the page90 minutes (locked)

    How does an alert reach the right place with enough context to act?

    You will be able to create an Alertmanager configuration under deploy/observability/alertmanager/.

    Fork the project repository to start working through the stages.

Phase 4Stages 9–100 of 2 stages verified

Phase 4 — Operate and learn

Run the incident hour deliberately and close the Beacon loop with usable runbooks and blameless postmortems.

  • Run the bad hour90 minutes (locked)

    What should the team do during the hour when Beacon is hurting users?

    You will be able to create an incident runbook for Beacon SLO burn under docs/runbooks/.

    Fork the project repository to start working through the stages.

  • Close the loop2 hours (locked)

    How does Beacon turn one bad night into a more reliable system?

    You will be able to finish the Beacon SLO-burn runbook with symptoms, confirmation, mitigation, escalation and rollback sections.

    Fork the project repository to start working through the stages.

About this path

Beacon exists to warn other people before endpoints fail silently, yet its own health is still mostly implied by deployment success. That is not operations. This course closes the loop: Beacon measures its own latency, errors, SLO burn and incident response with the same seriousness it applies to certificate expiry for others.

What you will learn: Prometheus instrumentation in Go, PromQL and histogram quantiles, SLI and SLO definition, Error budget burn-rate alerting, Incident response and postmortems. Build the project through cumulative challenges with beginner explanations and local verification.

Level
Complete beginner to independently building and operating the project
Format
10 cumulative stages. Every stage teaches the concept in full before any code, then gives you the thing to build and the run that proves it works
Before you start
You need Docker, kind, kubectl and Go 1.22 or newer. Prometheus, Grafana and Alertmanager all run locally in containers. The notification stage uses a free Discord webhook. There is zero cloud cost and no paid account requirement.
Official documentation (opens in a new tab)