SysDesignPrep.com
Study guide 19 of 183

Autoscaling and capacity planning

Horizontal versus vertical scaling, what to autoscale on, why autoscaling is too slow for spikes, warm pools and pre-scaling, serverless trade-offs, headroom, and how to plan capacity from numbers.

Reading is half of it. See this used in a real interview: walk through Design LeetCode →

"We will autoscale" is one of the most common phrases in system design answers and one of the least examined. Autoscaling handles gradual change well and sudden spikes badly, and some parts of a system cannot be autoscaled at all. Interviewers like candidates who know where it works, where it does not, and how to plan capacity with numbers.

Horizontal and vertical

  • Vertical scaling: a bigger machine. Simple, no code changes, and often the right first step for a database. Limits: the largest machine, a restart to resize, and one machine is still one failure domain.
  • Horizontal scaling: more machines behind a load balancer. Needs stateless services (sessions in a shared store, files in object storage), but scales far and survives the loss of any one machine. See scalability fundamentals.

Stateless tiers (web servers, API services, workers) scale horizontally and are what autoscalers manage. Stateful tiers (databases, caches, brokers) scale by sharding and replication, planned ahead, not by an autoscaler adding nodes in a panic.

What to scale on

SignalGood forWatch out for
CPU utilisationCPU-bound servicesI/O-bound services sit at low CPU while queueing
Requests per second per instanceweb and API tiersneeds a known capacity per instance
Queue depth or age of oldest messageworkersdepth alone hides slow jobs; age is better
Concurrent connectionsWebSocket and streaming gatewaysconnections are long-lived; scale-in must drain
Latencylast resortlatency rises for many reasons scaling will not fix
Custom (GPU utilisation, tokens per second)specialised fleetsmust reflect the true bottleneck

Target a utilisation that leaves headroom (often 50 to 70 % CPU), because new capacity takes minutes to arrive.

Autoscaling is slow

A typical reaction time: the metric must stay high for a minute or two, a new instance boots (seconds for containers, minutes for VMs), pulls an image, warms caches and passes health checks. Two to ten minutes from spike to capacity is normal.

That is fine for the daily curve, and useless for:

  • Flash crowds: a ticket sale, a push notification to millions, a viral post. Load arrives in seconds.
  • Scheduled spikes: a contest start, a product launch, a TV ad.

The answers:

  • Pre-scale for known events: scale up 30 minutes before the start, then let autoscaling handle the uncertainty. See Design LeetCode.
  • Warm pools of instances that are booted but idle, attached in seconds.
  • Headroom: run with enough spare capacity to absorb the fastest plausible surge for the minutes autoscaling needs.
  • Queues and load shedding to absorb or refuse the excess instead of collapsing. See rate limiting and resilience.
  • A waiting room for extreme flash crowds. See Design Ticketmaster.

Scaling in safely

Removing capacity is where outages happen:

  • Drain connections and in-flight requests before terminating an instance.
  • Cool-down periods and asymmetric policies (scale out fast, scale in slowly) prevent flapping.
  • Long-lived connections (WebSockets) must be moved gradually, or a scale-in event becomes a reconnect storm.
  • Scaling in a worker fleet must not kill jobs mid-way; workers finish or release their jobs first.

Serverless

Functions (Lambda, Cloud Functions, Vercel Functions) scale per request with no instances to manage and cost nothing when idle.

  • Good for: spiky or low-volume workloads, glue code, webhooks, scheduled tasks.
  • Costs: cold starts (milliseconds to seconds), per-invocation pricing that overtakes servers at steady high load, execution time limits, and concurrency limits that can throttle you during a spike.
  • Watch the downstream: a function that scales to 5,000 concurrent instances will open 5,000 database connections. Put a pooler or a queue in between.

Capacity planning with numbers

The back-of-envelope method from the estimation guide:

  1. Peak load: average × peak factor (2 to 3× for global consumer apps, more for event-driven ones).
  2. Per-instance capacity: measured by load testing, not guessed; for example 2,000 requests per second per instance at the target latency.
  3. Instances = peak ÷ per-instance capacity ÷ target utilisation. 30,000 rps ÷ 2,000 ÷ 0.6 ≈ 25 instances.
  4. Add N+1 or N+2 for failures and deploys, and per zone: if one of three zones fails, the other two must carry the load, so each zone runs at under two thirds of its capacity.
  5. Check the stateful tiers: databases, caches and brokers usually hit their limits before stateless services do.

Cost

Autoscaling saves money only if scale-in actually happens. Look at average utilisation: a fleet averaging 15 % CPU is over-provisioned. Use reserved or committed capacity for the baseline and on-demand or spot instances for the peaks; spot is cheap but can be reclaimed, so use it for workers and batch jobs that tolerate interruption.

In the interview

  • Say which tiers autoscale and on what signal.
  • Say how you handle spikes faster than autoscaling: pre-scaling, warm pools, headroom, queues.
  • Show one capacity calculation with utilisation and failure headroom.
  • Note what does not autoscale (the database) and how it is sized.

Checklist

  • Stateless tiers autoscale; stateful tiers are planned.
  • The right scaling signal per tier.
  • Headroom for the minutes autoscaling takes.
  • Pre-scaling for known events; queues and shedding for unknown ones.
  • Safe scale-in with draining.
  • A capacity estimate with failure headroom.

Open in your browser to sign in

Google does not allow sign-in inside this app's built-in browser. Open this page in Safari and sign in there. The link opens this same page.

Tap the ⋯ or share button at the top or bottom of the screen, then Open in browser. Or copy the link and paste it into Safari.