Project challenges / verified progress
Beacon: run it like an SRE

The engineering notebook

Spend the error budget

How does one reliability number become an engineering policy?

Loading statusStage 6 of 10

  • Workspace not ready
  • Agent not ready
Focus25:00
A small focus ritual

0 focus sessions completed. Every fourth session offers a longer break. Start each phase when you are ready.

Study time never unlocks verified lesson progress.

Loading...

Loading verified progress...

Loading GitHub account...
Phase 2 — Turn measurements into promises

Step 01 of 06

Learn the concept

One hundred percent sounds brave until you calculate the cost. It means no deploys, no dependency failures, no kernel bugs, no bad weather in someone else's data centre. SRE starts arguing from numbers instead of vibes.

THIRTY DAY ERROR BUDGETminutesGood time99.9 percent43156.8 minBudget0.1 percent43.2 minFast burn2 percent used0.864 min
A month has 43,200 minutes. At 99.9%, the bad allowance is 43.2 minutes. A fast-burn alert that fires after 2% of the budget in one hour is warning after less than a minute of monthly budget has disappeared.
Step 01

The ideas this is made of

SLI, SLO and SLA are not synonyms

The SLI is the measured ratio, such as 99.75% of checks completing within two seconds. The SLO is the internal target, such as 99.9% over 30 days. The SLA is a contract with customers and penalties. Mixing them causes bad conversations: engineering treats legal promises as dashboards, or sales treats internal goals as guarantees.

Each extra nine costs real work

A 99% monthly target allows 7.2 hours of badness. 99.9% allows 43.2 minutes. 99.99% allows 4.32 minutes. Moving one nine to the right means reducing allowed failure by a factor of ten. That may require redundancy, on-call staffing, safer deploys and fewer risky features. The number is a budget decision, not a tattoo.

Error budgets turn reliability into policy

If Beacon is comfortably within budget, teams can ship faster and accept measured risk. If the budget is burning too fast, the policy can pause risky deploys and focus on reliability work. This avoids the endless argument between feature speed and stability. Both sides agreed to the budget before the outage, which is the only civilized time to agree.

Rolling windows avoid calendar amnesia

A calendar-month SLO resets at midnight on the first, which can forgive an outage for no user-visible reason. A rolling 30-day window always asks what users experienced recently. It is harder to compute by hand and better for operations. Recording rules make that rolling math cheap and consistent.

The arithmetic nobody gets to skip
30 days * 24 hours * 60 minutes = 43,200 minutes

99.0% target  => 1.0% budget  => 432 minutes
99.9% target  => 0.1% budget  => 43.2 minutes
99.99% target => 0.01% budget => 4.32 minutes

A 20 minute outage spends:
20 / 43.2 = 46.3% of a 99.9% monthly budget

The difference between 99.9 and 99.99 is not a dot and a nine. It is whether a single 20 minute incident spends half the budget or more than four monthly budgets.

The reliability terms in one table

TermQuestionExample

SLI

What did we measure?

99.72% good checks

SLO

What target did we choose?

99.9% over 30d

SLA

What did we promise legally?

Credit below 99.5%

Error budget

How much badness remains?

18 minutes left

What these are called on the job

  • Error budget — The allowed amount of bad service implied by an SLO.

  • Burn rate — How quickly the service is consuming its budget relative to the allowed pace.

  • Recording rule — A PromQL expression evaluated on a schedule and stored as a new metric.