Serving LLMs on GPUs
The mechanics behind fast, affordable model serving: prefill and decode phases, the KV cache and paged attention, continuous batching, quantisation, tensor and pipeline parallelism, speculative decoding, prefix caching, routing and autoscaling GPU fleets, and the metrics that matter.
Reading is half of it. See this used in a real interview: walk through Design an LLM Inference Service →
Serving a large language model is a different problem from serving a typical web request. Each request runs for seconds, produces tokens one at a time, holds gigabytes of GPU memory while it runs, and the hardware costs dollars per hour. "Design an LLM inference service" is a newer interview question, and it rewards knowing why throughput and latency trade off the way they do on GPUs, and which techniques move the curve.
Two phases per request
- Prefill: the model processes the whole prompt in parallel and produces the first token. Compute-bound; its time grows with prompt length. It sets time to first token (TTFT).
- Decode: the model generates one token at a time, each step reading all model weights from GPU memory. Memory-bandwidth-bound; a single request uses the GPU poorly. It sets time per output token (TPOT) or inter-token latency.
Because one decode step for one request leaves most of the GPU idle, serving systems batch many requests' decode steps together.
The KV cache
During generation, the model stores keys and values for every previous token (the KV cache) so it does not recompute them. It grows with sequence length and batch size, and often limits how many requests fit on a GPU more than compute does. Memory per token depends on the model's layers and dimensions; long contexts multiply it.
Paged attention (popularised by vLLM) stores the KV cache in fixed-size blocks, like virtual memory pages, instead of one contiguous region per request. This removes fragmentation, lets more requests share the GPU, and allows sharing blocks between requests with a common prefix.
Continuous batching
Static batching waits for a batch to fill and for every request in it to finish, so short requests wait for long ones. Continuous (in-flight) batching adds new requests to the running batch at every decode step and removes finished ones immediately. GPU utilisation and throughput rise sharply, with lower average latency. Schedulers also decide how to mix prefill (bursty, compute-heavy) with ongoing decode steps, sometimes splitting long prefills into chunks, or running prefill and decode on separate GPU pools (disaggregated serving).
Making models cheaper
- Quantisation: store weights (and sometimes activations or the KV cache) in 8-bit or 4-bit formats instead of 16-bit. Less memory, more requests per GPU, faster decode; small quality loss that must be evaluated.
- Smaller or distilled models for easy requests, with a router choosing the model per request. See LLM systems.
- Speculative decoding: a small draft model proposes several tokens; the large model verifies them in one step. Same output distribution, fewer large-model steps.
- Prefix caching: reuse the KV cache of a shared prompt prefix (system prompts, few-shot examples, documents in a conversation) across requests, cutting prefill cost and TTFT. Route requests with the same prefix to the same replica to benefit.
Large models across GPUs
When a model does not fit on one GPU:
- Tensor parallelism: split each layer's matrices across GPUs in one server, which needs fast interconnect.
- Pipeline parallelism: put different layers on different GPUs or servers.
- Expert parallelism for mixture-of-experts models.
More GPUs per replica means fewer replicas for a fixed budget; capacity planning works in units of replicas. See queueing theory and capacity planning.
The serving system around the engine
- Gateway: authentication, per-tenant rate limits by tokens (not just requests), request validation. See rate limiting algorithms.
- Router: chooses a model and a replica, preferring replicas with free KV cache memory and cached prefixes; avoids overloading one replica.
- Queueing and priorities: interactive requests ahead of batch jobs; bounded queues with load shedding when full. See load shedding and backpressure.
- Streaming responses to clients with SSE, so users see tokens as they are generated. See WebSockets vs SSE vs long polling.
- Autoscaling on queue depth, KV cache utilisation and tokens per second, not CPU. Model loading takes minutes (weights are tens of gigabytes), so keep warm capacity and preload weights from fast local or cached storage. See autoscaling.
- Multi-region capacity: GPUs are scarce; route across regions when one is full.
Metrics
- TTFT and TPOT at p50 and p99, per model and prompt length bucket.
- Throughput in tokens per second per GPU, and cost per million tokens.
- Queue time, KV cache utilisation, batch size, prefix cache hit rate.
- Error and timeout rates, and quality metrics from evaluations.
The core trade-off: larger batches raise throughput and lower cost per token but increase per-token latency. Interactive products target a latency SLO and maximise throughput within it. See SLIs, SLOs and error budgets.
In the interview
For Design LLM Inference: explain prefill and decode, the KV cache as the main memory constraint, continuous batching with paged attention for utilisation, quantisation and speculative decoding for cost, prefix caching with prefix-aware routing, token-based rate limits and priorities at the gateway, SSE streaming, and autoscaling on queue depth with warm capacity because model loads are slow.
Checklist
- Prefill (TTFT) versus decode (TPOT) understood and measured separately.
- KV cache sized and managed with paged attention.
- Continuous batching; prefill and decode scheduling.
- Quantisation, routing to smaller models, speculative decoding.
- Prefix caching with prefix-aware routing.
- Token-based limits, priorities, bounded queues, streaming.
- Autoscaling on queue and cache metrics with warm, preloaded replicas.