Project challenges / verified progress
Beacon: run it like an SRE

The engineering notebook

See the running system

How do you know what a running system is doing before it wakes you up?

Loading statusStage 1 of 10

  • Workspace not ready
  • Agent not ready
Focus25:00
A small focus ritual

0 focus sessions completed. Every fourth session offers a longer break. Start each phase when you are ready.

Study time never unlocks verified lesson progress.

Loading...

Loading verified progress...

Loading GitHub account...
Phase 1 — Make Beacon visible

Step 01 of 06

Learn the concept

Beacon can now watch other people's endpoints, but no one is watching Beacon. That is not irony; it is technical debt with a pager. This stage gives names to the evidence an operator needs before choosing a tool.

ONE EVENT, THREE SIGNALSBeacon requestone check handledLogswhy this one failedrich, priceyMetricshow many and how fastcheap, aggregateTraceswhere time wentsampled path
The same request can leave three kinds of evidence. Logs preserve details, metrics preserve cheap totals, and traces preserve causality across components. Keeping those jobs separate prevents the classic mistake: trying to make Prometheus a log database with worse search.
Step 01

The ideas this is made of

Monitoring starts with questions you already know

Monitoring is the checklist: is the process up, are error rates rising, is disk full, are certificates expiring. Those questions are known before the incident. Observability is what helps when the symptom is new: a slow path only for one region, one tenant, and one release. Good operations needs both. Known alarms catch the common fires; exploratory evidence explains the weird smoke.

Metrics are cheap because they forget details

A counter such as beacon_checks_total{result="ok"} stores a number per label set, not a row per request. That is why Prometheus can hold months of service health on a laptop. The trade-off is real: you cannot ask which exact request failed unless you recorded that elsewhere. Metrics answer population questions: rate, ratio, percentile, saturation.

Cardinality is where innocent labels become outages

Suppose Beacon exposes ten useful time series. Add region with five values: now 50 series. Add endpoint with 4,000 targets: now 200,000. Add user_id with ten values during testing and it looks fine; with 100,000 users it becomes 20 billion potential series. Prometheus stores each as its own stream. A single unbounded label can take the server down.

Pull makes collection observable too

Prometheus usually scrapes targets over HTTP. If a scrape fails, Prometheus knows the target was unreachable and records up == 0. With push, the absence of data may mean the job ended, the network broke, or the push gateway is stale. Push still has a place for short-lived batch jobs. For long-running Beacon, pull is simpler and more honest.

Cardinality arithmetic on a shop checkout
metric: checkout_requests_total
labels: method=2, status=5, route=12
series: 2 * 5 * 12 = 120

add store_id=800
series: 120 * 800 = 96,000

add customer_id=2,000,000
series: 96,000 * 2,000,000 = 192,000,000,000

The metric name did not change. The code may have added one innocent label. The storage bill and query cost changed by nine orders of magnitude. That is why SREs ask about cardinality before they ask about dashboard colours.

The three signals do different work

SignalBest atCost trap

Logs

One event with context

High volume and retention

Metrics

Trends, ratios, alerts

High-cardinality labels

Traces

Cross-service latency paths

Sampling and storage cost

Profiles

CPU and memory hotspots

Usually not always-on

What these are called on the job

  • Time series — One stream of samples for one metric name and one exact set of labels.

  • Cardinality — The count of distinct label combinations a metric can produce.

  • Scrape — One Prometheus HTTP collection attempt against a target's metrics endpoint.

  • Known-unknown — A question you expected to ask, such as whether error rate is above 1%.