SysDesignPrep.com
Study guide 11 of 183

"Stateful vs stateless services"

Why stateless services scale and recover easily, where state has to live instead (databases, caches, sessions, object storage), and how to run genuinely stateful services such as WebSocket gateways, game servers, caches and stream processors with sticky routing, partitioning, replication and draining.

Reading is half of it. See this used in a real interview: walk through Design a URL Shortener →

"Make the service stateless" is standard advice in system design, and for good reason: stateless servers can be added, removed and replaced at will. But real systems are full of state: sessions, connections, caches, game worlds, stream processing windows. The skill is to know which parts of a design can be stateless, where their state should live instead, and how to run the parts that cannot be.

What stateless means

A stateless service keeps no data between requests that another instance would need. Every request carries or looks up what it needs (a token, a database row), so any instance can handle any request.

Benefits:

  • Horizontal scaling: add instances behind a load balancer. See scalability basics.
  • Simple failure handling: a crashed instance loses nothing; requests go elsewhere.
  • Easy deploys and autoscaling: instances are interchangeable. See autoscaling and deployment strategies.
  • Any load balancing algorithm works. See load balancing.

Where the state goes

Stateless services push state into systems designed to hold it:

StateWhere it lives
Business datadatabases
Sessionssigned tokens (JWT) or a shared session store such as Redis
Hot readsa distributed cache
Files and mediaobject storage
In-progress workqueues and workflow engines
Configurationa config service or feature flag system

Those stores are stateful and must be replicated, backed up and scaled, but there are few of them, and they are built for it.

Sessions: the classic example

If login state lives in one server's memory, the user must always return to that server (sticky sessions), and is logged out when it restarts. Fixes:

  • Signed tokens carry the session in each request; any server can verify them. Revocation needs short lifetimes or a deny list. See security and auth.
  • Shared session store (Redis) looked up per request, with replication.

Genuinely stateful services

Some components are stateful by nature, and making them stateless would be too slow or impossible:

Running stateful services well

  • Partition the state by key (user, match, document, cache key) so each instance owns a subset. See sharding and partitioning.
  • Route requests for a key to its owner: consistent hashing, a routing table or a registry. See consistent hashing.
  • Replicate or checkpoint so a crash does not lose everything: replicas for caches and databases, periodic checkpoints for stream processors, durable logs for anything that must survive.
  • Rebalance carefully when instances are added or removed, moving as little state as possible.
  • Drain before shutdown: stop accepting new work, hand off or finish existing work (ask clients to reconnect elsewhere, migrate sessions).
  • Recover by rebuilding state from durable storage (replaying a log, reloading a snapshot).

On Kubernetes, stateful workloads use stable identities and persistent volumes (StatefulSets) or, more often for core data, managed services.

Mostly stateless with soft state

Many services keep soft state: local caches that improve performance but can be lost safely. An instance with a warm cache is faster; a new one is slower until warmed, but correct. Consistent hashing of requests to instances improves hit rates while keeping instances replaceable.

In the interview

Point out which tiers are stateless (API servers behind a load balancer, scaled horizontally) and where state lives (database, cache, object storage, queues). For stateful components, say how they are partitioned, how requests are routed to the right instance, and what happens when one dies. For Design a URL Shortener, everything except the data stores is stateless. For Design WhatsApp, gateways hold connections, so explain the session registry, heartbeats and reconnects.

Checklist

  • Request-handling tiers stateless and interchangeable.
  • State in databases, caches, object storage, queues and token-based sessions.
  • No sticky sessions unless there is a real reason.
  • Stateful parts partitioned by key with deterministic routing.
  • Replication, checkpoints or durable logs for recovery.
  • Rebalancing and draining planned; soft state allowed but never required for correctness.

Open in your browser to sign in

Google does not allow sign-in inside this app's built-in browser. Open this page in Safari and sign in there. The link opens this same page.

Tap the ⋯ or share button at the top or bottom of the screen, then Open in browser. Or copy the link and paste it into Safari.