SysDesignPrep.com
Study guide 153 of 183

Load shedding and backpressure

Staying up when demand exceeds capacity: why overloaded systems collapse, backpressure in queues and streams, bounded buffers, load shedding by priority, adaptive concurrency limits, deadlines, brownouts and graceful degradation, and retry storms.

Reading is half of it. See this used in a real interview: walk through Design Ticketmaster →

Every system eventually receives more work than it can handle: a launch, a viral post, a failed region shifting traffic, a retry storm after an outage. Autoscaling helps, but it takes minutes, and some resources (a database primary) cannot scale quickly at all. What the system does in those minutes decides whether users see slightly degraded service or a total outage. Backpressure slows producers down; load shedding drops work deliberately. Both are signs of a mature design.

Why overload turns into collapse

Without protection, overload is self-reinforcing:

  1. Requests arrive faster than they are served; queues grow.
  2. Latency rises; clients time out, but the server keeps working on requests nobody is waiting for.
  3. Clients retry, adding more load.
  4. Memory fills with queued work; garbage collection and swapping slow everything further.
  5. Throughput of useful work drops toward zero, and the system stays down even after demand returns to normal.

Queueing theory explains the latency cliff near full utilisation. See queueing theory and capacity planning.

Backpressure

Backpressure propagates "slow down" upstream instead of buffering without limit:

  • Bounded queues: when full, producers block or get an error rather than adding more.
  • Pull-based consumption: consumers fetch work at their own pace (Kafka consumers, reactive streams) rather than having it pushed at them.
  • Flow control: TCP windows and HTTP/2 stream limits do this at the network level.
  • Explicit signals: return 429 Too Many Requests or 503 with Retry-After, so clients back off.

In streaming pipelines, backpressure lets a slow sink slow the whole pipeline instead of crashing a stage; the lag is visible and recovers later. See batch and stream processing.

Load shedding

When you cannot slow the source (users clicking, devices sending), reject some work early and cheaply so the rest succeeds:

  • Shed at the edge, before expensive work: the gateway or load balancer rejects when backends report overload.
  • Prioritise: keep checkout, login and paid traffic; shed analytics beacons, prefetches, background sync and bots first.
  • Reject new work, not in-flight work: finishing started requests is more valuable than starting new ones.
  • Prefer LIFO or deadline-aware queues under overload: the newest requests still have clients waiting; the oldest probably timed out.

Adaptive concurrency limits

Fixed limits are hard to set: too low wastes capacity, too high allows collapse. Adaptive limits (inspired by TCP congestion control) raise the allowed concurrency while latency stays near its baseline and cut it when latency rises. Requests beyond the limit are rejected immediately. Libraries implement this per service, so each protects itself without manual tuning.

Deadlines

Pass a deadline with every request (gRPC does this natively). Each service checks the remaining time before starting work and gives up if it cannot finish in time. Downstream calls inherit the shrinking budget. No work is wasted on requests the user has already abandoned. See tail latency.

Graceful degradation

Serve something useful with less work:

  • Return cached or slightly stale results instead of computing fresh ones. See caching.
  • Turn off expensive optional features (personalised ranking becomes popularity ranking; recommendations become a static list) with operational feature flags. See feature flags and A/B testing.
  • Reduce quality: lower video bitrate, fewer search results, smaller images.

Planned brownouts decide in advance which features degrade first.

Retry storms

Retries turn a small overload into a large one. Use:

  • Exponential backoff with jitter.
  • Retry budgets: retries limited to, say, 10 % of normal traffic.
  • Retry at one layer only.
  • Circuit breakers that stop calling a failing dependency for a while.

See rate limiting and resilience.

Queues for surges

For predictable spikes (a ticket sale opening), put users in a virtual waiting room and admit them at the rate the backend can handle, rather than letting everyone hit the system at once. Asynchronous work goes through queues that absorb bursts and are drained at a steady rate. See Design Ticketmaster and Design a Notification System.

In the interview

When asked "what happens if traffic spikes 10x?": autoscaling for the long term; for the first minutes, rate limits at the edge, priority-based shedding of non-critical traffic, adaptive concurrency limits per service, deadlines, cached and degraded responses, and backoff with retry budgets in clients. Mention a waiting room for planned spikes.

Checklist

  • Bounded queues everywhere; no unbounded buffering.
  • Backpressure via pull consumption and explicit 429 or 503 signals.
  • Priority-based shedding at the edge; finish in-flight work first.
  • Adaptive concurrency limits and propagated deadlines.
  • Planned degradation modes and operational flags.
  • Jittered backoff, retry budgets and circuit breakers.
  • Waiting rooms and queues for predictable surges.

Open in your browser to sign in

Google does not allow sign-in inside this app's built-in browser. Open this page in Safari and sign in there. The link opens this same page.

Tap the ⋯ or share button at the top or bottom of the screen, then Open in browser. Or copy the link and paste it into Safari.