SysDesignPrep.com
Study guide 21 of 183

Serverless architecture

When functions as a service fit a design and when they do not: event-driven triggers, scaling to zero and to thousands, cold starts, execution limits, state and connections, cost at low and high volume, managed building blocks, and common serverless patterns.

Reading is half of it. See this used in a real interview: walk through Design a URL Shortener →

Serverless platforms run your code in response to events (an HTTP request, a file upload, a queue message, a schedule) without you managing servers. They scale from zero to thousands of concurrent executions automatically and bill per invocation and compute time. That makes them excellent for some workloads and a poor fit for others. In interviews, mentioning serverless where it fits (and why not elsewhere) shows practical judgement about cost and operations.

How it works

  • You deploy a function with a handler and configure triggers.
  • The platform starts instances on demand, runs the handler, and keeps instances warm for a while for reuse.
  • Concurrency scales automatically with incoming events, up to account limits.
  • You pay for requests and compute time actually used, with nothing to pay when idle.

Modern platforms blur the line with containers (serverless containers that scale on request) and with longer-running, concurrent instances, but the model is the same: no servers to manage, scale tied to demand.

Where it shines

  • Spiky or unpredictable traffic: from nothing to a burst and back, without capacity planning. See autoscaling.
  • Event glue: resize an image when it lands in object storage; process a message from a queue; react to a database change. See media uploads and processing.
  • Scheduled tasks: nightly reports, cleanups. See delayed jobs and distributed cron.
  • Webhooks and lightweight APIs with modest traffic. See webhooks.
  • Parallel fan-out: split a big job into thousands of small invocations (thumbnailing, crawling pages, batch transforms). See Design a Web Crawler.
  • Small teams that want to minimise operations.

Limitations

  • Cold starts: a new instance must start the runtime and load code, adding from tens of milliseconds to seconds. Mitigate with small packages, fast runtimes, provisioned or pre-warmed concurrency, and keeping latency-critical paths warm.
  • Execution limits: maximum duration per invocation (minutes), memory and payload sizes. Long tasks must be split or moved to containers or workflows. See workflow orchestration.
  • Statelessness: no reliable local state between invocations; use external stores. See stateful vs stateless services.
  • Connection storms: thousands of concurrent instances each opening database connections can exhaust a relational database. Use connection proxies or poolers, or HTTP-based data APIs. See scaling a relational database.
  • Persistent connections like WebSockets need special support or a separate gateway service.
  • Debugging and local testing are harder; observability needs deliberate setup.
  • Vendor coupling to the platform's triggers and services.

Cost at scale

Serverless is cheap at low or bursty utilisation and can become expensive at high, steady load, where reserved containers or instances running at high utilisation cost less per request. A rough rule: if a workload would keep servers busy most of the day, compare the per-invocation bill with the cost of a right-sized fleet. See cost-aware system design.

Serverless building blocks

Beyond functions, many managed services are "serverless" in the same sense: object storage, managed queues, pub/sub, serverless databases and key-value stores, API gateways and workflow services. A whole system can be assembled from them with no servers to run, which suits startups and internal tools. See Kafka vs RabbitMQ vs SQS.

Patterns

  • Queue-triggered workers with retries and a dead-letter queue, idempotent handlers. See background jobs.
  • Fan-out and fan-in: a coordinator splits work into many invocations and a step collects results.
  • Edge functions for request rewriting, auth checks and personalisation close to users. See edge computing.
  • Strangler pattern: peel off features from a monolith into functions gradually.

Concurrency control

Unlimited scaling can overwhelm downstream systems. Set concurrency limits per function, put queues between functions and fragile dependencies, and apply backpressure. See load shedding and backpressure.

In the interview

Use serverless for the parts it fits and say why: "Thumbnail generation runs as functions triggered by uploads, because volume is bursty and each task is short; the redirect service runs on always-on containers, because traffic is steady and latency-critical, so cold starts and per-request pricing would hurt." For Design a Notification System, queue-triggered functions can be a simple sender tier if volumes are moderate. For Design a URL Shortener, the analytics pipeline is a natural serverless candidate.

Checklist

  • Serverless for spiky, event-driven, short tasks and glue code.
  • Containers or servers for steady, latency-critical, long-running or connection-heavy work.
  • Cold start mitigation on user-facing paths.
  • External state, pooled database access, idempotent handlers.
  • Concurrency limits and queues to protect dependencies.
  • Cost compared at expected utilisation.

Open in your browser to sign in

Google does not allow sign-in inside this app's built-in browser. Open this page in Safari and sign in there. The link opens this same page.

Tap the ⋯ or share button at the top or bottom of the screen, then Open in browser. Or copy the link and paste it into Safari.