SysDesignPrep.com
Study guide 13 of 183

Networking for system design

DNS, TCP and UDP, HTTP/1.1 versus HTTP/2 and HTTP/3, TLS, proxies and reverse proxies, and the latency numbers behind them, at the depth a system design interview expects.

Reading is half of it. See this used in a real interview: walk through Design YouTube →

Most system design answers assume the network works and costs nothing. Interviewers notice when a candidate knows what actually happens between a click and a response: how a name becomes an address, how many round trips a new connection costs, and why the answer to "how do we make this faster for users in Sydney" is usually about round trips, not servers. This guide covers the parts that change designs.

What happens when you open a URL

  1. DNS lookup. The browser asks a recursive resolver (usually the ISP's or a public one such as 1.1.1.1) for the name's address. If nothing is cached, the resolver walks from the root servers to the top-level domain to the domain's authoritative servers. Cached answers return in a few milliseconds; a cold lookup can take 50 to 200 ms.
  2. TCP handshake. SYN, SYN-ACK, ACK: one round trip before any data can flow.
  3. TLS handshake. TLS 1.3 adds one more round trip (zero on resumption); TLS 1.2 adds two.
  4. HTTP request and response. One more round trip, plus server time, plus transfer time for the body.

From Sydney to Virginia a round trip is about 200 ms, so a cold HTTPS request costs roughly 600 ms before the server has done any work. That single number explains CDNs, connection reuse and edge termination better than any diagram.

DNS as a design tool

DNS is not just name lookup; it is the first load balancer.

  • TTL decides how quickly a change propagates. Low TTLs (30 to 60 s) allow fast failover but increase resolver traffic; high TTLs (hours) are cheap but slow to change. Many clients and resolvers ignore very low TTLs, so never rely on DNS alone for failover in seconds.
  • Geo DNS and latency-based routing return a nearby region's address based on the resolver's location. Imprecise when users use a distant public resolver; EDNS Client Subnet helps.
  • Weighted records shift a percentage of traffic, which is how many canary and regional migrations start.

In an interview, say "DNS points users at the nearest region with a short TTL, and a global load balancer or anycast handles fast failover", which shows you know DNS is slow to change.

TCP versus UDP

TCP gives a reliable, ordered byte stream with congestion control. The costs are the handshake and head-of-line blocking: if one packet is lost, everything behind it waits for the retransmission, even data for unrelated requests sharing the connection.

UDP sends independent datagrams with no handshake, ordering or retransmission. Applications that prefer late data to be dropped rather than delay everything use it: voice and video calls, game state updates, DNS queries, and QUIC (which rebuilds reliability on top).

UseProtocolWhy
Web APIs, databases, most servicesTCPcorrectness matters more than a few milliseconds
Live video calls, multiplayer stateUDP (often via WebRTC)a late frame is useless; drop it and move on
Modern web trafficQUIC (HTTP/3)TCP-like reliability without head-of-line blocking or the extra handshake

HTTP/1.1, HTTP/2 and HTTP/3

  • HTTP/1.1 allows one request at a time per connection, so browsers open about six connections per host. Keep-alive reuses connections across requests.
  • HTTP/2 multiplexes many requests over one TCP connection with header compression. Far fewer connections, but a lost TCP packet still stalls every stream on it.
  • HTTP/3 runs over QUIC on UDP. Streams are independent, so one loss stalls only its stream; connection setup combines transport and TLS into one round trip (zero on resumption); and a connection survives a phone switching from Wi-Fi to cellular because it is identified by an id, not by IP and port.

For mobile users on lossy networks, HTTP/3 measurably cuts tail latency. Behind your load balancer, HTTP/2 or gRPC between services is the norm.

TLS and where to terminate it

TLS encrypts traffic and proves the server's identity with a certificate. The design question is where it ends:

  • At the CDN or edge: the handshake happens close to the user (a 20 ms round trip instead of 200 ms). The edge then talks to your origin over a separate, often long-lived, encrypted connection.
  • At the load balancer: simplest; traffic inside the data centre is plain or re-encrypted.
  • End to end to each service (mutual TLS): required in zero-trust networks; a service mesh usually does it so applications do not.

Terminating at the edge is the single biggest latency win for global users after caching.

Proxies, reverse proxies and gateways

A forward proxy acts for clients (a corporate egress proxy). A reverse proxy acts for servers: clients think they are talking to the service, and the proxy terminates TLS, routes, caches, compresses, rate limits and shields backends. Nginx, Envoy and HAProxy are reverse proxies; an API gateway is a reverse proxy with product features (authentication, quotas, request transformation); a load balancer is a reverse proxy whose main job is spreading load. In diagrams, one box called "API gateway / load balancer" is usually enough unless the interview is about one of them. See load balancing.

Connections are expensive; reuse them

Every new connection costs handshakes, memory and, for databases, a backend process. Patterns that come up constantly:

  • Keep-alive and connection pools between services, sized to the number of concurrent requests, not the request rate.
  • A pooler in front of Postgres (PgBouncer), because thousands of application instances each opening tens of connections overwhelms the database.
  • Long-lived connections for push (WebSockets, SSE) change capacity planning from requests per second to concurrent connections; see real-time systems.

Bandwidth and latency numbers

PathRound trip
Same machine (loopback)< 0.1 ms
Same data centre~0.5 ms
Same region, different zone1–2 ms
Across a continent50–100 ms
Across the world150–300 ms

A server NIC is 10 to 100 Gbit/s; a typical mobile connection is 5 to 50 Mbit/s with high variance. Transfer time matters for large bodies: a 5 MB response over 10 Mbit/s takes 4 seconds no matter how fast your servers are, which is why compression, image variants and pagination are latency features. More numbers in the estimation guide.

In the interview

Bring networking up when it changes the design:

  • Users far from the servers: count round trips, then terminate TLS and cache at the edge.
  • Real-time or mobile clients: persistent connections, reconnection with backoff, and HTTP/3 or WebSockets.
  • Service-to-service: gRPC over HTTP/2 with connection pooling and timeouts on every call.
  • Media and large files: transfer time dominates, so send less (compression, variants, ranges).

Checklist

  • Round trips for a cold request from the farthest user, and how you cut them.
  • DNS TTLs and what actually does fast failover.
  • Where TLS terminates.
  • TCP or UDP for each stream, and why.
  • Connection reuse and pooling, especially in front of the database.
  • Payload sizes on mobile networks.

Open in your browser to sign in

Google does not allow sign-in inside this app's built-in browser. Open this page in Safari and sign in there. The link opens this same page.

Tap the ⋯ or share button at the top or bottom of the screen, then Open in browser. Or copy the link and paste it into Safari.