SysDesignPrep.com
Study guide 136 of 183

Presence and managing millions of connections

Holding millions of persistent WebSocket connections and showing who is online: connection gateways, routing messages to the right server, heartbeats, presence with TTLs, fan-out of status changes, typing indicators, reconnect storms and deploys.

Reading is half of it. See this used in a real interview: walk through Design WhatsApp →

Chat apps, collaborative editors and live dashboards keep a persistent connection open to every active client. At tens of millions of users, that means tens of millions of open sockets spread across hundreds of servers, and the system must know which server holds each user so a message can reach them. On top of that, "online" dots and "typing..." indicators generate surprising amounts of traffic. These are the details behind "how does the message actually get to Bob's phone?".

Connection gateways

Dedicated gateway servers hold client connections (WebSocket, or MQTT on mobile) and do little else:

  • Terminate TLS, authenticate the connection once, then keep it open.
  • Forward client messages to backend services; deliver server messages to clients.
  • Stay mostly stateless beyond the open connections, so they can be added and drained freely.

An efficient event-driven server can hold hundreds of thousands to over a million mostly idle connections, limited by memory per connection, file descriptors and heartbeat traffic. Keep business logic out of gateways so they rarely need deploys. See real-time systems.

Finding the user's connection

When Alice sends Bob a message, the backend must know which gateway holds Bob's connection (or connections, one per device):

  • A session registry maps user id to (gateway id, connection id), written on connect and removed on disconnect, usually in Redis or a similar fast store with TTLs.
  • The message service looks up Bob's gateways and sends the message to each one through an internal channel (direct RPC, or a per-gateway queue or pub/sub topic).
  • If Bob has no live connection, the message is stored and a push notification is sent. See push notifications.

Alternatively, gateways subscribe to topics (per user or per channel) on a pub/sub layer, and publishers do not need to know gateway ids. This is simpler for group channels and large fan-outs. See Design Slack.

Heartbeats and dead connections

Mobile connections often die silently (a network change, a phone going to sleep). So:

  • Client and server exchange heartbeats every 30 to 60 seconds; missing a few marks the connection dead.
  • Choose intervals with battery and NAT timeouts in mind: too frequent drains batteries, too rare lets mobile carriers drop idle connections.
  • Remove registry entries when a connection is declared dead, and give entries a TTL that heartbeats refresh, so crashed gateways do not leave ghosts.

Presence

"Online" and "last seen" are derived from connections:

  • A user is online if any device has a live connection (or has been active within a short window).
  • Store presence with a TTL refreshed by heartbeats; expiry means offline, even if the disconnect was never observed.
  • Debounce flapping: a phone switching from Wi-Fi to cellular should not broadcast offline then online.

Fan-out of presence changes

Presence is cheap to store and expensive to broadcast. If each user has 500 contacts and status changes several times an hour, broadcasting every change to every contact is enormous. Limit it:

  • Send presence only to contacts who are currently online and looking (subscribed to that conversation or contact list view).
  • Fetch presence on demand when opening a chat, instead of pushing continuously.
  • For large groups and channels, show approximate counts ("1,204 online") rather than per-member status.
  • Batch and rate-limit updates.

Typing indicators are similar: ephemeral, sent only to the open conversation, rate-limited, never stored.

Reconnect storms and deploys

If a gateway crashes, all its clients reconnect at once; if a region fails, millions do. Protect the system:

  • Clients reconnect with exponential backoff and jitter.
  • Session resumption: the client sends the last message id it saw, and the server sends only what was missed, from the message store, not from in-memory buffers.
  • Rate-limit connection setup and authentication.
  • For deploys, drain gateways gradually: stop accepting new connections, ask clients to reconnect elsewhere over minutes, then restart.

Load balancing long connections

Round-robin on connect leaves old servers full and new ones empty after scaling, because connections last hours. Balance on connection count, and rebalance gently by asking some clients to reconnect. Layer 4 load balancers or DNS spread connections across gateways. See load balancing.

In the interview

For Design WhatsApp: gateways holding persistent connections, a session registry mapping users to gateways, messages stored durably then routed to the recipient's gateways (or push if offline), acknowledgements for delivery receipts, heartbeats with TTL-based presence, limited presence fan-out, and backoff with session resumption on reconnect.

Checklist

  • Dedicated, thin connection gateways sized for connections, not CPU.
  • Session registry with TTLs, or pub/sub topics, to route messages.
  • Durable message store as the source of truth; push when offline.
  • Heartbeats tuned for battery and NAT; TTL-based presence.
  • Presence and typing sent only to interested, online viewers.
  • Backoff, jitter, session resumption and graceful draining.

Open in your browser to sign in

Google does not allow sign-in inside this app's built-in browser. Open this page in Safari and sign in there. The link opens this same page.

Tap the ⋯ or share button at the top or bottom of the screen, then Open in browser. Or copy the link and paste it into Safari.