Project challenges / verified progress
Beacon: run it like an SRE

The engineering notebook

Instrument the service

What should Beacon expose so Prometheus can measure Beacon itself?

Loading statusStage 2 of 10

  • Workspace not ready
  • Agent not ready
Focus25:00
A small focus ritual

0 focus sessions completed. Every fourth session offers a longer break. Start each phase when you are ready.

Study time never unlocks verified lesson progress.

Loading...

Loading verified progress...

Loading GitHub account...
Phase 1 — Make Beacon visible

Step 01 of 06

Learn the concept

A service without metrics is a black box with logs taped to the side. Beacon already measures latency and expiry for other endpoints; now it must expose its own request rate, errors and duration. The loop starts to close here.

METRIC TYPES BY PROMISESummaryclient-side quantilesrare choiceHistogrambucketed observationsSLO readyGaugevalue can move both waysin flightCountermonotonic event countrate later
Each type is a promise to future queries. A counter promises monotonicity, so `rate()` can repair restarts. A histogram promises fixed buckets, so Prometheus can aggregate latency across replicas. Pick the promise that matches the phenomenon.
Step 01

The ideas this is made of

A counter is a count of events, not a current value

Requests handled, checks completed and errors seen are counters. They only increase until the process restarts. Prometheus expects that reset and rate() turns the staircase into per-second speed. If you put request totals in a gauge, a restart looks like negative traffic and an alert may either vanish or fire for nonsense.

A gauge is for measurements that can reverse

Gauges fit in-flight checks, queue depth, open connections and the Unix timestamp of the last successful scrape. They can rise and fall without violating their meaning. A gauge is not a flexible counter. If the real-world quantity cannot decrease except by process restart, use a counter and let Prometheus handle the reset correctly.

Histograms are useful because buckets aggregate

A histogram stores cumulative bucket counters such as le="0.25" seconds. Because those are counters, Prometheus can add them across pods and then estimate a percentile with histogram_quantile. The price is that buckets are chosen up front. If your largest bucket is one second, a five-second outage becomes +Inf and the detail is gone.

Names carry units so dashboards do not guess

Prometheus names should read like facts: beacon_http_requests_total, beacon_check_duration_seconds, beacon_checks_in_flight. Use base units, not milliseconds or mixed units, and reserve suffixes for their meaning. A name ending _total should be a counter. A duration in seconds should end _seconds. Boring names save incident minutes.

A tiny handler with real metrics
package main

import (
	"log"
	"net/http"
	"time"

	"github.com/prometheus/client_golang/prometheus"
	"github.com/prometheus/client_golang/prometheus/promhttp"
)

var requests = prometheus.NewCounterVec(
	prometheus.CounterOpts{Name: "demo_requests_total", Help: "Requests by code."},
	[]string{"code"},
)

var duration = prometheus.NewHistogram(prometheus.HistogramOpts{
	Name:    "demo_request_duration_seconds",
	Help:    "Request duration in seconds.",
	Buckets: []float64{0.05, 0.1, 0.25, 0.5, 1},
})

func main() {
	prometheus.MustRegister(requests, duration)
	http.HandleFunc("/", func(w http.ResponseWriter, r *http.Request) {
		start := time.Now()
		w.WriteHeader(http.StatusOK)
		requests.WithLabelValues("200").Inc()
		duration.Observe(time.Since(start).Seconds())
	})
	http.Handle("/metrics", promhttp.Handler())
	log.Fatal(http.ListenAndServe(":8080", nil))
}

The labels are bounded: code has a small vocabulary. The duration is a histogram, not a summary, because later stages need SLO math across every Beacon replica. The example is small enough to compile and large enough to show the real library.

Which metric type fits the fact?

TypeUse it forDo not use it for

Counter

Requests, errors, checks

Current queue depth

Gauge

In-flight work, config size

Lifetime totals

Histogram

Latency and sizes

Unique users

Summary

Process-local quantiles

Fleet-wide SLO math

What these are called on the job

  • Collector — A value registered with the Prometheus client that can produce metrics during a scrape.

  • Bucket — A histogram boundary; samples less than or equal to it increment that cumulative counter.

  • Base unit — Seconds, bytes and ratios rather than milliseconds, megabytes or percent.