SysDesignPrep.com
Study guide 14 of 183

Global traffic management, DNS and anycast

How users reach the nearest healthy region: DNS resolution and TTLs, GeoDNS and latency-based routing, anycast, global load balancers, health checks and failover, the limits of DNS-based failover, and steering traffic during incidents and migrations.

Reading is half of it. See this used in a real interview: walk through Design a URL Shortener →

A global product runs in several regions, and every user should reach a nearby, healthy one. The first decision happens before any request reaches your servers: when the client resolves your domain name and opens a connection. DNS, anycast and global load balancers are the tools, and each has quirks that matter during an outage. This guide covers the layer above the regional load balancer.

DNS basics that matter

  • The client asks a recursive resolver (from the ISP, or a public one), which asks the authoritative servers for your domain and caches the answer for its TTL.
  • A short TTL (30 to 60 seconds) lets you change answers quickly; a long TTL reduces query load and latency.
  • Resolvers and clients do not always respect TTLs; some cache longer. Expect a long tail of traffic going to old answers for minutes after a change.
  • The authoritative server sees the resolver's location, not the user's; the EDNS client subnet extension passes part of the client address to improve accuracy.

See networking fundamentals.

Routing policies

PolicyHow it picksUse
GeoDNSanswer by the resolver's (or client subnet's) locationsend Europe to the EU region, for latency or data residency
Latency-basedanswer with the region that has the lowest measured latency from that networkbest performance
Weightedsplit by percentagesgradual migrations, canarying a new region
Failoverprimary answer unless its health checks failactive-passive setups

Managed DNS services combine these with health checks of each region's endpoints.

Anycast

With anycast, the same IP address is announced from many locations; internet routing (BGP) delivers each packet to the topologically nearest one. Benefits:

  • No DNS caching delay: if a site withdraws its announcement, traffic shifts to the next nearest within seconds to a minute.
  • Natural absorption of DDoS traffic across many sites.
  • One address worldwide, simpler for clients.

CDNs and public DNS resolvers rely on anycast. Long-lived TCP connections can break if routes shift mid-connection, so anycast works best in front of edge proxies that terminate connections close to users. See CDN and edge.

Global load balancers

Cloud global load balancers combine anycast front ends with proxies at the edge: the user connects to the nearest edge location, which forwards to the best healthy backend region over the provider's network. You get fast failover without DNS changes, plus TLS termination close to users. The trade-off is dependence on one provider's edge.

Health checks and failover

  • Probe each region from many locations, checking a deep health endpoint (can it serve real requests?) rather than just "is the port open".
  • Require several consecutive failures before failing over, and several successes before failing back, to avoid flapping.
  • Make sure the remaining regions can absorb the traffic; failover that overloads the survivors turns one regional outage into a global one. See queueing theory and capacity planning.
  • Practise failover regularly. See testing distributed systems.

Limits of DNS failover

DNS failover takes at least the TTL plus resolver caching, often several minutes for a tail of users, and clients with open connections keep using the failed region until those connections break. For faster failover use anycast or a global load balancer, and make clients retry against alternative endpoints.

Steering for other reasons

  • Data residency: route users to the region that must hold their data, based on account, not location, since users travel. See privacy and data deletion.
  • Sticky regions: users whose data lives in one region (home region) should be routed there, or the edge should forward to it.
  • Draining a region for maintenance with weighted routing, gradually.
  • Load shedding across regions when one is overloaded.

See multi-region architecture.

Mobile and API clients

Apps can do their own steering: fetch a list of endpoints from a config service, measure latency, and fall back to another region on errors. Large streaming and messaging apps use client-side steering to reach the best edge or cache server. See Design Netflix and Design WhatsApp.

In the interview

One or two sentences usually suffice: "Users resolve to the nearest healthy region via latency-based DNS with 60-second TTLs (or anycast through a global load balancer for faster failover); health checks remove a failing region, and the others have headroom to absorb its traffic." For a read-heavy global service like Design a URL Shortener, add edge caching of redirects.

Checklist

  • DNS TTLs chosen with failover speed in mind; stale caches expected.
  • Geo or latency-based routing; weighted routing for migrations.
  • Anycast or a global load balancer for fast failover.
  • Deep health checks with hysteresis; capacity to absorb a failed region.
  • Residency and home-region routing based on the account.
  • Client-side fallback for apps.

Open in your browser to sign in

Google does not allow sign-in inside this app's built-in browser. Open this page in Safari and sign in there. The link opens this same page.

Tap the ⋯ or share button at the top or bottom of the screen, then Open in browser. Or copy the link and paste it into Safari.