Step 01 of 06
Learn the concept
One hundred percent sounds brave until you calculate the cost. It means no deploys, no dependency failures, no kernel bugs, no bad weather in someone else's data centre. SRE starts arguing from numbers instead of vibes.
The ideas this is made of
SLI, SLO and SLA are not synonyms
The SLI is the measured ratio, such as 99.75% of checks completing within two seconds. The SLO is the internal target, such as 99.9% over 30 days. The SLA is a contract with customers and penalties. Mixing them causes bad conversations: engineering treats legal promises as dashboards, or sales treats internal goals as guarantees.
Each extra nine costs real work
A 99% monthly target allows 7.2 hours of badness. 99.9% allows 43.2 minutes. 99.99% allows 4.32 minutes. Moving one nine to the right means reducing allowed failure by a factor of ten. That may require redundancy, on-call staffing, safer deploys and fewer risky features. The number is a budget decision, not a tattoo.
Error budgets turn reliability into policy
If Beacon is comfortably within budget, teams can ship faster and accept measured risk. If the budget is burning too fast, the policy can pause risky deploys and focus on reliability work. This avoids the endless argument between feature speed and stability. Both sides agreed to the budget before the outage, which is the only civilized time to agree.
Rolling windows avoid calendar amnesia
A calendar-month SLO resets at midnight on the first, which can forgive an outage for no user-visible reason. A rolling 30-day window always asks what users experienced recently. It is harder to compute by hand and better for operations. Recording rules make that rolling math cheap and consistent.
30 days * 24 hours * 60 minutes = 43,200 minutes
99.0% target => 1.0% budget => 432 minutes
99.9% target => 0.1% budget => 43.2 minutes
99.99% target => 0.01% budget => 4.32 minutes
A 20 minute outage spends:
20 / 43.2 = 46.3% of a 99.9% monthly budgetThe difference between 99.9 and 99.99 is not a dot and a nine. It is whether a single 20 minute incident spends half the budget or more than four monthly budgets.
The reliability terms in one table
| Term | Question | Example |
|---|---|---|
SLI | What did we measure? | 99.72% good checks |
SLO | What target did we choose? | 99.9% over 30d |
SLA | What did we promise legally? | Credit below 99.5% |
Error budget | How much badness remains? | 18 minutes left |
What these are called on the job
Error budget — The allowed amount of bad service implied by an SLO.
Burn rate — How quickly the service is consuming its budget relative to the allowed pace.
Recording rule — A PromQL expression evaluated on a schedule and stored as a new metric.
