SysDesignPrep.com
Study guide 150 of 183

The thundering herd problem

When many clients act at the same moment and overwhelm a system: cache stampedes, retry storms, reconnect storms after outages, synchronised cron jobs and notification opens, cold starts after deploys, and the fixes: jitter, request coalescing, early refresh, backoff, admission control and staggering.

Reading is half of it. See this used in a real interview: walk through Design a Distributed Cache →

A system that handles its average load comfortably can still fall over when thousands of clients do the same thing at the same instant. The thundering herd problem appears in many forms: a popular cache key expires, a service recovers and every client reconnects, every server runs its cron job at midnight, a push notification makes millions of people open the app together. Interviewers like it because it tests whether you think about time, not just volume.

Cache stampede

A hot key expires; every request misses at once and hits the database with the same expensive query.

Fixes:

  • Request coalescing (single flight): the first miss fetches; concurrent requests for the same key wait for that result.
  • Stale-while-revalidate: serve the expired value while one background request refreshes it.
  • Early probabilistic refresh: as expiry approaches, each request has a small, growing chance to refresh early, so one refresh happens before the key expires.
  • Jittered TTLs: avoid many keys expiring at the same moment.
  • Locks or leases on recomputation so only one worker rebuilds an expensive entry.

See cache invalidation and Design a Distributed Cache.

Retry storms

A dependency slows down; clients time out and retry, multiplying load exactly when the dependency is weakest. Layers that each retry multiply further.

Fixes: exponential backoff with jitter, retry budgets (retries capped as a share of traffic), retry at one layer only, and circuit breakers that stop calling a failing dependency. See rate limiting and resilience.

Reconnect storms

A gateway restarts or a region recovers; millions of persistent connections reconnect at once, each requiring TLS handshakes, authentication and state sync.

Fixes:

  • Clients reconnect after a randomised delay with exponential backoff.
  • Session resumption so reconnecting clients fetch only what they missed.
  • Admission control: limit new connections per second per gateway; reject politely with a retry hint.
  • Drain servers gradually during deploys instead of restarting all at once.

See presence and connection management and Design WhatsApp.

Synchronised schedules

Cron jobs at 0 0 * * * on thousands of machines, mobile apps syncing "every hour on the hour", daily reports kicked off at midnight UTC: all start together.

Fixes: add jitter to schedules (spread over minutes), use a central scheduler that rate-limits dispatch, and avoid round times for background syncs. See delayed jobs and distributed cron.

Human herds

Cold starts

After a deploy, a scale-up or a cache flush, new instances have empty caches and cold connection pools, and they all send their misses to the database at once. Warm caches before taking traffic, ramp traffic to new instances gradually, and avoid flushing entire cache clusters. See autoscaling.

Lock and queue herds

Many workers waiting on the same lock or queue are woken together when it becomes available, but only one can proceed; the rest spin and go back to sleep. Wake one waiter at a time, use fair queues, or partition work so fewer workers contend for the same item. See distributed locks and leases.

The common fixes

FixHerd it breaks
Jitter (randomised timing)synchronised expiry, retries, schedules, reconnects
Coalescing and single flightduplicate work for the same key
Serve stale, refresh in backgroundcache expiry spikes
Backoff and retry budgetsretry storms
Admission control and waiting roomsbursts of new sessions and users
Staggering and gradual rampsnotifications, deploys, scale-ups
Pre-warming and pre-scalingpredictable events

In the interview

When your design has a cache, persistent connections, scheduled work or big launches, add one sentence: "TTLs are jittered and misses are coalesced; clients reconnect with jittered backoff and resume sessions; scheduled jobs are spread over a window; the sale opens through a waiting room." It shows you think about synchronised behaviour, not just averages.

Checklist

  • Identify moments when many clients act together.
  • Jitter on TTLs, retries, reconnects and schedules.
  • Request coalescing and stale-while-revalidate for hot keys.
  • Backoff, retry budgets and circuit breakers.
  • Admission control, waiting rooms and staggered sends.
  • Warm caches and ramp traffic after deploys and scale-ups.

Open in your browser to sign in

Google does not allow sign-in inside this app's built-in browser. Open this page in Safari and sign in there. The link opens this same page.

Tap the ⋯ or share button at the top or bottom of the screen, then Open in browser. Or copy the link and paste it into Safari.