Project challenges / verified progress
Beacon: run it like an SRE

The engineering notebook

Ask the data questions

How does Prometheus collect Beacon metrics and turn samples into answers?

Loading statusStage 3 of 10

  • Workspace not ready
  • Agent not ready
Focus25:00
A small focus ritual

0 focus sessions completed. Every fourth session offers a longer break. Start each phase when you are ready.

Study time never unlocks verified lesson progress.

Loading...

Loading verified progress...

Loading GitHub account...
Phase 1 — Make Beacon visible

Step 01 of 06

Learn the concept

A /metrics endpoint is only a diary page until something reads it every few seconds. Prometheus turns those pages into time series, then PromQL asks the operational questions. The language looks odd because time itself is part of the type system.

SCRAPE STORE QUERY ALERT01ScrapeGET /metrics on a clock15s here02Storename plus labelsseries03QueryPromQL over windowsrate, sum04Alertrules evaluate oftensymptoms
Prometheus is both collector and time-series database. The scrape interval decides how often evidence arrives; PromQL decides what question is asked over that evidence. Alerts are just PromQL queries with memory and consequences.
Step 01

The ideas this is made of

The label set is the identity of a series

beacon_checks_total{result="success"} and beacon_checks_total{result="failure"} are different series. Add instance="beacon:8080" and they are different again. Prometheus does not store a table with a result column; it stores many independent streams. That model makes aggregation powerful and cardinality dangerous for the same reason.

`rate()` turns counter samples into speed

Counters are staircases. rate(beacon_checks_total[5m]) looks at the last five minutes, handles resets, and returns per-second average increase. Without rate(), a counter graph mostly tells you the process has been alive for a while. With rate(), it tells you traffic, error rate and burn speed. A gauge already has a current value, so its rate is rarely the question.

Aggregation decides what detail survives

sum by (result) keeps the result label and discards instance. sum without (instance) says the same thing from the other direction. If you aggregate too early, you lose the dimension needed to debug. If you aggregate too late, the query is noisy or expensive. Incident queries should preserve the label that points to an owner or a failing slice.

Histograms need the `le` label until the last moment

Prometheus histograms expose bucket counters labelled by le, meaning less than or equal. A percentile query must sum bucket rates while preserving le, then pass the result to histogram_quantile. Drop le in the sum by and the function has no bucket boundaries left. That beginner error produces empty results or nonsense.

Three honest queries
sum(rate(http_requests_total[5m])) by (code)

sum(rate(http_requests_total{code=~"5.."}[5m]))
/
sum(rate(http_requests_total[5m]))

histogram_quantile(
  0.95,
  sum(rate(http_request_duration_seconds_bucket[5m])) by (le)
)

The first query keeps a debugging label. The second computes an error ratio from two counter rates. The third preserves le until histogram_quantile can use the bucket boundaries.

PromQL values beginners confuse

ValueWhat it isTypical use

Instant vector

Series at one evaluation time

Current error ratio

Range vector

Series samples over a window

Input to rate()

Scalar

One number

Threshold like 0.99

String

Rare metadata value

Almost never in alerts

What these are called on the job

  • Target — One scrape destination, such as beacon:8080.

  • Job — A group of targets sharing a scrape configuration.

  • Range selector — The [5m] part that gives PromQL a window of samples.

  • Recording rule — A PromQL expression evaluated and stored as a new series.