SLIs, SLOs and error budgets
Defining reliability with numbers: choosing SLIs, setting SLO targets, availability nines and what they allow, error budgets and how teams spend them, burn-rate alerting, and SLAs versus SLOs.
Reading is half of it. See this used in a real interview: walk through Design a Metrics and Monitoring System →
"Highly available" means nothing until it has a number. Site reliability practice turns reliability into measurable targets, and it gives teams a rule for when to ship faster and when to stop and fix things. In interviews, stating non-functional requirements as SLOs, and knowing what 99.99 % actually allows, is a quick way to sound like someone who has run production systems.
The vocabulary
- SLI (service level indicator): a measurement of user-visible behaviour, as a ratio of good events to total events. "The proportion of checkout requests that succeed in under 500 ms."
- SLO (service level objective): the target for that SLI over a window. "99.9 % over 28 days."
- SLA (service level agreement): a contract with customers, with consequences (credits) if broken. Set SLAs looser than SLOs, so you notice and fix problems before they cost money.
- Error budget: 100 % minus the SLO. With a 99.9 % SLO, 0.1 % of requests may fail; that budget can be "spent" on incidents, risky deploys and experiments.
Choosing good SLIs
Measure what users experience, as close to them as practical:
| Kind of service | SLIs |
|---|---|
| Request/response API | availability (non-5xx share), latency (share under a threshold) |
| Data pipeline | freshness (share of data processed within N minutes), correctness |
| Storage | durability, read and write availability |
| Streaming or push | delivery latency, share delivered |
Prefer "share of requests faster than 300 ms" to "p99 latency", because ratios add up cleanly over time and across instances. Measure at the load balancer or client where possible; server-side metrics miss requests that never arrived.
What the nines allow
| SLO | Downtime per 30 days | Per year |
|---|---|---|
| 99 % | 7.2 hours | 3.65 days |
| 99.9 % | 43 minutes | 8.8 hours |
| 99.95 % | 22 minutes | 4.4 hours |
| 99.99 % | 4.3 minutes | 53 minutes |
| 99.999 % | 26 seconds | 5.3 minutes |
Two consequences worth saying out loud:
- Each nine costs roughly ten times more in engineering and infrastructure. 99.99 % leaves no room for a manual response to an incident: detection and failover must be automatic.
- Dependencies multiply: a service calling five dependencies at 99.9 % each cannot itself promise more than about 99.5 % unless it degrades gracefully when they fail.
Not everything needs the same target: login and checkout may need 99.95 %, an internal reporting page 99 %.
Error budgets in practice
The budget turns reliability into a shared decision rather than an argument:
- Budget remaining: ship features, run experiments, take reasonable risks.
- Budget exhausted: freeze risky launches, prioritise reliability work until the SLI recovers.
It also makes clear that 100 % is the wrong goal: a perfectly reliable service ships too slowly, and users cannot tell the difference between 99.99 % and 100 % when their own network is less reliable than that.
Alerting on burn rate
Alerting on every error spike wakes people for nothing; alerting on monthly SLO breach is too late. Burn-rate alerts fire when the budget is being consumed too fast:
- A fast burn (for example, 2 % of the 28-day budget in one hour, a rate that would exhaust it in about two days) pages someone immediately.
- A slow burn (for example, 10 % in three days) opens a ticket.
- Each alert checks a long and a short window together, so it fires quickly and stops quickly when the problem ends.
This is how the monitoring system design turns metrics into useful pages.
Using SLOs in a design
- State the important targets as SLOs in the non-functional requirements: "redirects: 99.99 % availability, p99 under 50 ms" (see Design a URL Shortener).
- Let them drive architecture: 99.99 % implies multi-zone, automatic failover and no single region dependency; 99.9 % may allow a simpler setup.
- Plan graceful degradation so dependencies failing do not count against the SLO (cached responses, partial results). See rate limiting and resilience.
Checklist
- SLIs as good/total ratios, measured near users.
- SLO targets per user journey, not one number for everything.
- The downtime each target allows, and what it requires (automation, redundancy).
- Error budget policy: what happens when it runs out.
- Burn-rate alerts with fast and slow windows.
- SLAs looser than SLOs.