Step 01 of 06
Learn the concept
A server can be healthy while the product is broken. CPU can be calm while every certificate warning is late. An SLI is the discipline of measuring the thing users would complain about, not the thing the host finds convenient.
The ideas this is made of
A good SLI follows the user's pain
For Beacon, users care that scheduled checks run, complete quickly and report accurate failures before certificates expire. They do not care that the server has spare CPU during a missed check. An SLI should move when user trust moves. If a metric can be green while users complain, it is at best diagnostic and at worst a distraction.
Event-based ratios avoid vague availability
Instead of saying Beacon is available, count events. A valid event is a scheduled check that should have run. A good event is one that completed successfully within, say, two seconds and produced a result. If 99,700 of 100,000 valid events are good, the SLI is 99.7%. That number can be argued with and improved.
Specification and implementation must be separate
The SLI specification might say: the proportion of scheduled checks that produce a correct result within two seconds. The implementation might use beacon_checks_total and beacon_check_duration_seconds_bucket. Keeping those separate lets you improve instrumentation without changing the product promise. It also exposes gaps where the current metric only approximates the desired truth.
Server-side measurement is often too kind
A server can report success after it writes a response, while the client timed out waiting for it. A queue can accept work and return 202 while the check never runs. Measuring at the server is convenient, but the best SLI is as close as practical to the user's experience. For Beacon, the scheduler completion point is more honest than the API accept point.
A photo service uploads 50,000 images in one day.
Valid events: uploads whose file type and size are allowed = 48,000
Good events: valid uploads processed and visible within 30 seconds = 47,520
SLI: 47,520 / 48,000 = 0.99 = 99%Invalid file types are excluded because users did not ask for a valid service action. Slow processing is counted as bad because the product promise includes visibility, not just accepting bytes.
Comfort metrics versus user SLIs
| Metric | What it says | Why it may lie |
|---|---|---|
CPU usage | Host is busy or idle | Product can fail at 10% CPU |
HTTP 200 rate | API accepted requests | Work may fail later |
Check success ratio | Beacon completed checks | Closer to user value |
Check p95 latency | Checks finish in time | Needs good buckets |
What these are called on the job
SLI — Service Level Indicator: a measurement of service quality.
Valid event — An event that should count in the SLI denominator.
Good event — A valid event that met the quality threshold.
Measurement point — Where in the system the evidence is collected.
