Project challenges / verified progress
Beacon: run it like an SRE

The engineering notebook

Measure what users feel

Which measurement tells whether Beacon is actually working for its users?

Loading statusStage 5 of 10

  • Workspace not ready
  • Agent not ready
Focus25:00
A small focus ritual

0 focus sessions completed. Every fourth session offers a longer break. Start each phase when you are ready.

Study time never unlocks verified lesson progress.

Loading...

Loading verified progress...

Loading GitHub account...
Phase 2 — Turn measurements into promises

Step 01 of 06

Learn the concept

A server can be healthy while the product is broken. CPU can be calm while every certificate warning is late. An SLI is the discipline of measuring the thing users would complain about, not the thing the host finds convenient.

EVENTS TO USER TRUTH01Valid eventschecks that should countdenominator02Good eventssuccessful and fastnumerator03Ratiogood divided by validSLI04Decisionship, pause, or pagepolicy
The denominator is as important as the numerator. Count synthetic test traffic, cancelled requests or invalid input carelessly and the SLI becomes a comfortable lie. The ratio should track what a user would agree was good service.
Step 01

The ideas this is made of

A good SLI follows the user's pain

For Beacon, users care that scheduled checks run, complete quickly and report accurate failures before certificates expire. They do not care that the server has spare CPU during a missed check. An SLI should move when user trust moves. If a metric can be green while users complain, it is at best diagnostic and at worst a distraction.

Event-based ratios avoid vague availability

Instead of saying Beacon is available, count events. A valid event is a scheduled check that should have run. A good event is one that completed successfully within, say, two seconds and produced a result. If 99,700 of 100,000 valid events are good, the SLI is 99.7%. That number can be argued with and improved.

Specification and implementation must be separate

The SLI specification might say: the proportion of scheduled checks that produce a correct result within two seconds. The implementation might use beacon_checks_total and beacon_check_duration_seconds_bucket. Keeping those separate lets you improve instrumentation without changing the product promise. It also exposes gaps where the current metric only approximates the desired truth.

Server-side measurement is often too kind

A server can report success after it writes a response, while the client timed out waiting for it. A queue can accept work and return 202 while the check never runs. Measuring at the server is convenient, but the best SLI is as close as practical to the user's experience. For Beacon, the scheduler completion point is more honest than the API accept point.

A valid-event ratio outside Beacon
A photo service uploads 50,000 images in one day.

Valid events: uploads whose file type and size are allowed = 48,000
Good events: valid uploads processed and visible within 30 seconds = 47,520
SLI: 47,520 / 48,000 = 0.99 = 99%

Invalid file types are excluded because users did not ask for a valid service action. Slow processing is counted as bad because the product promise includes visibility, not just accepting bytes.

Comfort metrics versus user SLIs

MetricWhat it saysWhy it may lie

CPU usage

Host is busy or idle

Product can fail at 10% CPU

HTTP 200 rate

API accepted requests

Work may fail later

Check success ratio

Beacon completed checks

Closer to user value

Check p95 latency

Checks finish in time

Needs good buckets

What these are called on the job

  • SLI — Service Level Indicator: a measurement of service quality.

  • Valid event — An event that should count in the SLI denominator.

  • Good event — A valid event that met the quality threshold.

  • Measurement point — Where in the system the evidence is collected.