Global traffic management, DNS and anycast
How users reach the nearest healthy region: DNS resolution and TTLs, GeoDNS and latency-based routing, anycast, global load balancers, health checks and failover, the limits of DNS-based failover, and steering traffic during incidents and migrations.
Reading is half of it. See this used in a real interview: walk through Design a URL Shortener →
A global product runs in several regions, and every user should reach a nearby, healthy one. The first decision happens before any request reaches your servers: when the client resolves your domain name and opens a connection. DNS, anycast and global load balancers are the tools, and each has quirks that matter during an outage. This guide covers the layer above the regional load balancer.
DNS basics that matter
- The client asks a recursive resolver (from the ISP, or a public one), which asks the authoritative servers for your domain and caches the answer for its TTL.
- A short TTL (30 to 60 seconds) lets you change answers quickly; a long TTL reduces query load and latency.
- Resolvers and clients do not always respect TTLs; some cache longer. Expect a long tail of traffic going to old answers for minutes after a change.
- The authoritative server sees the resolver's location, not the user's; the EDNS client subnet extension passes part of the client address to improve accuracy.
Routing policies
| Policy | How it picks | Use |
|---|---|---|
| GeoDNS | answer by the resolver's (or client subnet's) location | send Europe to the EU region, for latency or data residency |
| Latency-based | answer with the region that has the lowest measured latency from that network | best performance |
| Weighted | split by percentages | gradual migrations, canarying a new region |
| Failover | primary answer unless its health checks fail | active-passive setups |
Managed DNS services combine these with health checks of each region's endpoints.
Anycast
With anycast, the same IP address is announced from many locations; internet routing (BGP) delivers each packet to the topologically nearest one. Benefits:
- No DNS caching delay: if a site withdraws its announcement, traffic shifts to the next nearest within seconds to a minute.
- Natural absorption of DDoS traffic across many sites.
- One address worldwide, simpler for clients.
CDNs and public DNS resolvers rely on anycast. Long-lived TCP connections can break if routes shift mid-connection, so anycast works best in front of edge proxies that terminate connections close to users. See CDN and edge.
Global load balancers
Cloud global load balancers combine anycast front ends with proxies at the edge: the user connects to the nearest edge location, which forwards to the best healthy backend region over the provider's network. You get fast failover without DNS changes, plus TLS termination close to users. The trade-off is dependence on one provider's edge.
Health checks and failover
- Probe each region from many locations, checking a deep health endpoint (can it serve real requests?) rather than just "is the port open".
- Require several consecutive failures before failing over, and several successes before failing back, to avoid flapping.
- Make sure the remaining regions can absorb the traffic; failover that overloads the survivors turns one regional outage into a global one. See queueing theory and capacity planning.
- Practise failover regularly. See testing distributed systems.
Limits of DNS failover
DNS failover takes at least the TTL plus resolver caching, often several minutes for a tail of users, and clients with open connections keep using the failed region until those connections break. For faster failover use anycast or a global load balancer, and make clients retry against alternative endpoints.
Steering for other reasons
- Data residency: route users to the region that must hold their data, based on account, not location, since users travel. See privacy and data deletion.
- Sticky regions: users whose data lives in one region (home region) should be routed there, or the edge should forward to it.
- Draining a region for maintenance with weighted routing, gradually.
- Load shedding across regions when one is overloaded.
See multi-region architecture.
Mobile and API clients
Apps can do their own steering: fetch a list of endpoints from a config service, measure latency, and fall back to another region on errors. Large streaming and messaging apps use client-side steering to reach the best edge or cache server. See Design Netflix and Design WhatsApp.
In the interview
One or two sentences usually suffice: "Users resolve to the nearest healthy region via latency-based DNS with 60-second TTLs (or anycast through a global load balancer for faster failover); health checks remove a failing region, and the others have headroom to absorb its traffic." For a read-heavy global service like Design a URL Shortener, add edge caching of redirects.
Checklist
- DNS TTLs chosen with failover speed in mind; stale caches expected.
- Geo or latency-based routing; weighted routing for migrations.
- Anycast or a global load balancer for fast failover.
- Deep health checks with hysteresis; capacity to absorb a failed region.
- Residency and home-region routing based on the account.
- Client-side fallback for apps.