SysDesignPrep.com
Study guide 174 of 183

SLIs, SLOs and error budgets

Defining reliability with numbers: choosing SLIs, setting SLO targets, availability nines and what they allow, error budgets and how teams spend them, burn-rate alerting, and SLAs versus SLOs.

Reading is half of it. See this used in a real interview: walk through Design a Metrics and Monitoring System →

"Highly available" means nothing until it has a number. Site reliability practice turns reliability into measurable targets, and it gives teams a rule for when to ship faster and when to stop and fix things. In interviews, stating non-functional requirements as SLOs, and knowing what 99.99 % actually allows, is a quick way to sound like someone who has run production systems.

The vocabulary

  • SLI (service level indicator): a measurement of user-visible behaviour, as a ratio of good events to total events. "The proportion of checkout requests that succeed in under 500 ms."
  • SLO (service level objective): the target for that SLI over a window. "99.9 % over 28 days."
  • SLA (service level agreement): a contract with customers, with consequences (credits) if broken. Set SLAs looser than SLOs, so you notice and fix problems before they cost money.
  • Error budget: 100 % minus the SLO. With a 99.9 % SLO, 0.1 % of requests may fail; that budget can be "spent" on incidents, risky deploys and experiments.

Choosing good SLIs

Measure what users experience, as close to them as practical:

Kind of serviceSLIs
Request/response APIavailability (non-5xx share), latency (share under a threshold)
Data pipelinefreshness (share of data processed within N minutes), correctness
Storagedurability, read and write availability
Streaming or pushdelivery latency, share delivered

Prefer "share of requests faster than 300 ms" to "p99 latency", because ratios add up cleanly over time and across instances. Measure at the load balancer or client where possible; server-side metrics miss requests that never arrived.

What the nines allow

SLODowntime per 30 daysPer year
99 %7.2 hours3.65 days
99.9 %43 minutes8.8 hours
99.95 %22 minutes4.4 hours
99.99 %4.3 minutes53 minutes
99.999 %26 seconds5.3 minutes

Two consequences worth saying out loud:

  • Each nine costs roughly ten times more in engineering and infrastructure. 99.99 % leaves no room for a manual response to an incident: detection and failover must be automatic.
  • Dependencies multiply: a service calling five dependencies at 99.9 % each cannot itself promise more than about 99.5 % unless it degrades gracefully when they fail.

Not everything needs the same target: login and checkout may need 99.95 %, an internal reporting page 99 %.

Error budgets in practice

The budget turns reliability into a shared decision rather than an argument:

  • Budget remaining: ship features, run experiments, take reasonable risks.
  • Budget exhausted: freeze risky launches, prioritise reliability work until the SLI recovers.

It also makes clear that 100 % is the wrong goal: a perfectly reliable service ships too slowly, and users cannot tell the difference between 99.99 % and 100 % when their own network is less reliable than that.

Alerting on burn rate

Alerting on every error spike wakes people for nothing; alerting on monthly SLO breach is too late. Burn-rate alerts fire when the budget is being consumed too fast:

  • A fast burn (for example, 2 % of the 28-day budget in one hour, a rate that would exhaust it in about two days) pages someone immediately.
  • A slow burn (for example, 10 % in three days) opens a ticket.
  • Each alert checks a long and a short window together, so it fires quickly and stops quickly when the problem ends.

This is how the monitoring system design turns metrics into useful pages.

Using SLOs in a design

  • State the important targets as SLOs in the non-functional requirements: "redirects: 99.99 % availability, p99 under 50 ms" (see Design a URL Shortener).
  • Let them drive architecture: 99.99 % implies multi-zone, automatic failover and no single region dependency; 99.9 % may allow a simpler setup.
  • Plan graceful degradation so dependencies failing do not count against the SLO (cached responses, partial results). See rate limiting and resilience.

Checklist

  • SLIs as good/total ratios, measured near users.
  • SLO targets per user journey, not one number for everything.
  • The downtime each target allows, and what it requires (automation, redundancy).
  • Error budget policy: what happens when it runs out.
  • Burn-rate alerts with fast and slow windows.
  • SLAs looser than SLOs.

Open in your browser to sign in

Google does not allow sign-in inside this app's built-in browser. Open this page in Safari and sign in there. The link opens this same page.

Tap the ⋯ or share button at the top or bottom of the screen, then Open in browser. Or copy the link and paste it into Safari.