SysDesignPrep.com
Study guide 181 of 183

Building LLM features: serving, latency and cost

System design for LLM-powered products: streaming responses, the cost and latency of tokens, batching and KV caches in inference, prompt caching, rate limits and fallbacks, guardrails, evaluation, and when to use RAG, tools or fine-tuning.

Reading is half of it. See this used in a real interview: walk through Design an LLM Inference Service →

More and more interview questions include an LLM: a support assistant, a summariser, a coding helper, a search answer box. The model is usually a given (an API or a self-hosted open model); the design is everything around it. LLM calls are slow, expensive, rate-limited and non-deterministic, which changes how you build the rest of the system. This guide covers the decisions that come up.

The shape of an LLM call

  • Latency has two parts: time to first token (prefill: reading the prompt, hundreds of milliseconds to a few seconds for long prompts) and time per output token (decoding: tens of milliseconds per token). A 500-token answer can take 10 to 20 seconds to finish.
  • Cost is per token, input and output priced separately, with output usually several times more expensive. Long prompts (big documents, long conversations) dominate cost even when answers are short.
  • Limits: providers enforce requests per minute and tokens per minute per account or model, and outages happen.

Back-of-envelope: 1 M requests a day × 3,000 input tokens × $3 per million input tokens is about $9,000 a day before output tokens. Estimate it early; it often reshapes the design.

Stream the response

Never make a user wait 15 seconds for a complete answer. Stream tokens to the client as they are generated (server-sent events or WebSockets) so text starts appearing after the first token. Design the backend to proxy the stream, handle client disconnects (cancel the generation to stop paying for it), and store the final response when it completes. See real-time systems.

Spending fewer tokens

  • Prompt caching: providers cache a repeated prompt prefix (system prompt, tools, a large document) and charge a fraction for cached tokens while cutting time to first token. Put stable content first and variable content last.
  • Response caching: cache answers to identical (or, carefully, semantically similar) requests for deterministic tasks.
  • Smaller models for easier work: route classification, extraction and short summaries to a small, fast model; reserve the largest model for the hard requests.
  • Retrieve, do not stuff: send the relevant passages rather than whole documents. See vector search and RAG.
  • Trim conversation history with summaries of older turns.

Self-hosted inference

When you run open models yourself, the serving system matters as much as the model:

  • Continuous batching: GPUs are efficient only when processing many sequences at once; continuous batching adds and removes requests from the running batch every step instead of waiting for a whole batch to finish.
  • KV cache: each in-flight sequence keeps attention state in GPU memory; this, not compute, usually limits concurrency. Paged KV caches (vLLM) and prefix sharing let more requests fit.
  • Admission control: when the GPUs are full, queue or shed requests rather than letting latency explode for everyone.
  • Capacity is in tokens per second, not requests per second, and GPUs are expensive and slow to provision, so autoscaling is coarse and headroom matters. See Design an LLM Inference Service and autoscaling.

Reliability around the model

  • Timeouts sized for long generations, and cancellation when the user leaves.
  • Retries on rate limits and transient errors, with backoff, and an idempotency key for side effects triggered by the model.
  • Fallbacks: a second provider or model when the first is down or over its limits, and a non-LLM degraded mode where possible.
  • Per-user and per-tenant quotas in tokens, enforced before calling the model, so one user cannot spend the whole budget. See rate limiting and resilience.
  • Queue non-interactive work (bulk summarisation, document processing) as background jobs with their own rate limits. See background jobs.

Structured output and tools

When the model must produce something a program consumes (JSON, a SQL filter, a function call), use the provider’s structured output or tool-calling features and validate the result against a schema, retrying or falling back on failure. For agents that call tools, treat every tool call as an untrusted request: check permissions, cap the number of steps, and log everything.

Safety and guardrails

  • Prompt injection: text from users, documents or web pages can contain instructions. Never let retrieved content grant permissions; enforce access control in your code, not in the prompt.
  • Data protection: decide what user data may be sent to a provider, redact where needed, and check the provider’s retention terms.
  • Output checks: moderation for user-facing text, and grounding checks (does the answer cite retrieved sources?) where hallucination is costly.

Evaluation and observability

LLM output changes with every prompt edit or model version, so treat prompts like code:

  • A test set of real inputs with expected properties, scored automatically (exact checks, rubric grading by a model, human review for a sample) and run before every change.
  • Production logging of prompts, responses, latency, tokens and cost per feature, with user feedback (thumbs up or down) linked to traces.
  • Dashboards for time to first token, total latency, error and rate-limit rates, and cost per request and per user.

RAG, tools or fine-tuning?

  • Prompting with retrieval (RAG) when the model needs knowledge it does not have, especially private or changing data.
  • Tools when it needs to act or look things up live (search, a database, an API).
  • Fine-tuning to change style, format or behaviour on a narrow task, or to make a smaller model perform like a larger one; not the first tool for adding knowledge.

Checklist

  • Latency budget split into time to first token and generation; streaming to the client.
  • Token cost estimate per request and per day.
  • Prompt caching, model routing and retrieval to cut tokens.
  • Timeouts, retries, fallbacks and per-tenant token quotas.
  • Structured output validated against a schema.
  • Prompt-injection defences and data-handling rules.
  • An evaluation set and production monitoring of quality, latency and cost.

Open in your browser to sign in

Google does not allow sign-in inside this app's built-in browser. Open this page in Safari and sign in there. The link opens this same page.

Tap the ⋯ or share button at the top or bottom of the screen, then Open in browser. Or copy the link and paste it into Safari.