Step 01 of 06
Learn the concept
Beacon can now watch other people's endpoints, but no one is watching Beacon. That is not irony; it is technical debt with a pager. This stage gives names to the evidence an operator needs before choosing a tool.
The ideas this is made of
Monitoring starts with questions you already know
Monitoring is the checklist: is the process up, are error rates rising, is disk full, are certificates expiring. Those questions are known before the incident. Observability is what helps when the symptom is new: a slow path only for one region, one tenant, and one release. Good operations needs both. Known alarms catch the common fires; exploratory evidence explains the weird smoke.
Metrics are cheap because they forget details
A counter such as beacon_checks_total{result="ok"} stores a number per label set, not a row per request. That is why Prometheus can hold months of service health on a laptop. The trade-off is real: you cannot ask which exact request failed unless you recorded that elsewhere. Metrics answer population questions: rate, ratio, percentile, saturation.
Cardinality is where innocent labels become outages
Suppose Beacon exposes ten useful time series. Add region with five values: now 50 series. Add endpoint with 4,000 targets: now 200,000. Add user_id with ten values during testing and it looks fine; with 100,000 users it becomes 20 billion potential series. Prometheus stores each as its own stream. A single unbounded label can take the server down.
Pull makes collection observable too
Prometheus usually scrapes targets over HTTP. If a scrape fails, Prometheus knows the target was unreachable and records up == 0. With push, the absence of data may mean the job ended, the network broke, or the push gateway is stale. Push still has a place for short-lived batch jobs. For long-running Beacon, pull is simpler and more honest.
metric: checkout_requests_total
labels: method=2, status=5, route=12
series: 2 * 5 * 12 = 120
add store_id=800
series: 120 * 800 = 96,000
add customer_id=2,000,000
series: 96,000 * 2,000,000 = 192,000,000,000The metric name did not change. The code may have added one innocent label. The storage bill and query cost changed by nine orders of magnitude. That is why SREs ask about cardinality before they ask about dashboard colours.
The three signals do different work
| Signal | Best at | Cost trap |
|---|---|---|
Logs | One event with context | High volume and retention |
Metrics | Trends, ratios, alerts | High-cardinality labels |
Traces | Cross-service latency paths | Sampling and storage cost |
Profiles | CPU and memory hotspots | Usually not always-on |
What these are called on the job
Time series — One stream of samples for one metric name and one exact set of labels.
Cardinality — The count of distinct label combinations a metric can produce.
Scrape — One Prometheus HTTP collection attempt against a target's metrics endpoint.
Known-unknown — A question you expected to ask, such as whether error rate is above 1%.
