Queueing theory and capacity planning
The few queueing ideas every system designer should know: Little's law, utilisation and why latency explodes near 100 %, headroom, concurrency limits, sizing thread pools, connection pools and worker fleets, and planning capacity for peaks and failures.
Reading is half of it. See this used in a real interview: walk through Design a Rate Limiter →
Every server, database and worker pool is a queue: requests arrive, wait, get served and leave. A little queueing theory explains why a system that is fine at 70 % utilisation falls over at 95 %, how many workers or connections you need, and how much headroom to plan for. It turns capacity questions in interviews from guesses into short, convincing calculations.
Little's law
For any stable system:
L = λ × W
- L: average number of items in the system (in progress plus waiting).
- λ: arrival rate (requests per second).
- W: average time each spends in the system.
It works for anything. Examples:
- A service handles 2,000 requests per second with 50 ms average latency, so about 2,000 × 0.05 = 100 requests are in flight at any moment. That is the concurrency you need (threads, connections, async slots).
- A queue receives 500 jobs per second and each job waits and runs for 4 seconds in total, so about 2,000 jobs are in the system.
- A database pool: 1,000 queries per second at 5 ms each needs about 5 connections busy on average; size for peaks and variance, perhaps 15 to 20, not 500.
Utilisation and latency
Utilisation is the fraction of time a server is busy: ρ = λ / μ, where μ is the service rate. As utilisation approaches 100 %, waiting time does not grow linearly; it explodes. In the simplest model (random arrivals, one server), average time in the system is the service time divided by (1 − ρ):
| Utilisation | Time in system relative to service time |
|---|---|
| 50 % | 2× |
| 70 % | 3.3× |
| 80 % | 5× |
| 90 % | 10× |
| 95 % | 20× |
| 99 % | 100× |
Real systems differ in detail, but the shape always holds. That is why services target roughly 50 to 70 % utilisation at peak, and why a small traffic increase near saturation causes a sudden latency cliff and timeouts. Variance makes it worse: bursty arrivals and variable service times increase queueing at any utilisation. See tail latency.
More servers, shared queue
Several servers pulling from one shared queue wait far less than the same servers each with their own queue, because no server sits idle while another has a backlog. This is why a central work queue with a pool of workers beats random assignment to per-worker queues, and why "join the shortest queue" or "least requests" load balancing beats round robin under uneven load. See load balancing.
Overload and bounded queues
When arrivals exceed capacity, an unbounded queue grows forever: latency rises until every request times out, and the system does useless work for clients who have already left. Protect yourself:
- Bound queues and reject (or shed) when full; a fast "try again later" beats a slow failure.
- Timeouts and deadlines propagated downstream, so work for abandoned requests is dropped.
- Concurrency limits per service, ideally adaptive (increase while latency is fine, back off when it rises).
- Prioritise important work when shedding.
See rate limiting and resilience.
Capacity planning
- Measure or estimate per-instance capacity at an acceptable latency (from a load test), not at the breaking point.
- Forecast peak demand: daily peak, seasonal peaks, launches and events (ticket sales spike orders of magnitude above normal). See Design Ticketmaster.
- Add headroom for utilisation targets and variance.
- Plan for failure: with N+1 or N+2 redundancy, or enough capacity to lose a whole zone (with three zones, each must handle half the load). See multi-region architecture.
- Account for scaling speed: if new capacity takes minutes, you need buffer for the minutes before it arrives. See autoscaling.
Example: peak 30,000 requests per second, one instance handles 1,000 at good latency, target 60 % utilisation, so 30,000 / 600 = 50 instances; spread across three zones that must survive the loss of one, about 75.
Worker fleets
For background workers and GPU servers, Little's law sizes the fleet: if jobs arrive at 200 per second and take 3 seconds each, you need at least 600 concurrent worker slots, plus headroom. Track queue depth and age as the scaling signal, since CPU can look fine while the backlog grows. See Design a Job Scheduler and Design LLM Inference.
Checklist
- Little's law to size concurrency, pools and fleets.
- Peak utilisation targets of roughly 50 to 70 %, lower for latency-critical paths.
- Shared queues and least-loaded balancing.
- Bounded queues, deadlines, concurrency limits and load shedding.
- Capacity for peak plus headroom plus losing a zone, with scaling delay covered.