Multi-region architecture
Why and when to run in several regions, active-passive versus active-active, routing users to a region, replicating data across regions, conflict handling, failover and evacuation, and data residency rules.
Reading is half of it. See this used in a real interview: walk through Design Netflix →
"What happens if a whole region goes down?" is one of the most common late-interview questions, and "multi-region" is one of the most expensive answers to give without thinking. Running in several regions buys lower latency for distant users and survival of a regional outage, and it costs consistency, money and complexity. This guide is about choosing the right amount of it.
Why go multi-region
- Latency: users in Sydney should not pay 200 ms round trips to Virginia for every request.
- Availability: survive the loss of an entire region (rare, but it happens, and some customers require it).
- Data residency: laws or contracts may require some users’ data to stay in a country or bloc (the EU, India).
Be clear which of these drives the design, because they lead to different architectures. Latency alone is often solved by a CDN and edge caching without multi-region data. See CDNs and edge computing.
The main shapes
| Shape | Writes go to | Failover | Consistency | Cost and complexity |
|---|---|---|---|---|
| Single region, multi-zone | one region | zone failures only | simple, strong | lowest |
| Active-passive (warm standby) | the primary region | promote the standby; minutes | strong in primary; standby lags | moderate; standby capacity is idle |
| Active-active, home region per user | each user’s home region | move users to another region | strong per user; cross-user data eventual | high |
| Active-active, multi-writer | any region | instant; regions are peers | eventual with conflict resolution, or slow global consensus | highest |
Most companies should start with multi-zone in one region (zones are separate data centres with independent power, a millisecond apart), which already survives the most common failures. Add regions when one of the three reasons demands it.
Routing users to a region
- Geo DNS or latency-based DNS sends users to the nearest healthy region; slow to change because of DNS caching.
- Anycast with a global load balancer (Cloudflare, AWS Global Accelerator, Google’s load balancer) routes to the nearest region and fails over in seconds.
- Home-region routing: the user’s id or tenant maps to a home region, and the edge forwards their writes there, even if they are travelling.
See networking for system design for why DNS alone is not fast failover.
Replicating data across regions
Cross-region round trips are 50 to 150 ms, so the core choice is whether a write waits for other regions.
- Asynchronous replication: the write commits locally and is shipped to other regions within a second or so. Fast; on regional failure the last moments of writes may be lost (RPO of seconds).
- Synchronous or consensus replication (Spanner, CockroachDB, a Raft group spanning regions): the write waits for a majority of regions. No data loss on failover; every write pays cross-region latency.
Pick per data type, as the CAP and PACELC guide explains: balances and inventory may justify synchronous replication; feeds, profiles and analytics do not.
Conflicts in active-active
If two regions can accept writes to the same record, they will sometimes do so concurrently. Options, best first:
- Avoid conflicts by design: give each record a home region (a user’s data lives in their home region; other regions forward writes there). Most "active-active" systems at large companies are really this.
- Merge automatically: CRDTs for counters, sets and text; append-only data that never conflicts.
- Last-writer-wins by timestamp: simple, but silently drops concurrent writes; acceptable only where that is harmless.
- Surface conflicts to the application or user to resolve.
Failover and evacuation
- Detect: health checks from several outside vantage points, not just from inside the failing region.
- Decide: automate failover for stateless tiers; for databases, automatic promotion risks split brain, so many teams require a human decision or use consensus-based stores that handle it.
- Shift traffic gradually and watch the receiving regions, which must have headroom to absorb the load (running each region at under 60 % or so).
- Fail back deliberately once the region is healthy and caught up.
The most important practice is rehearsal: evacuate a region on purpose, regularly. Netflix’s region evacuations are the well-known example. A failover path that has never been exercised usually fails when needed.
Know your targets: RTO (how long until service is restored) and RPO (how much recent data may be lost). Asynchronous replication with automated failover typically gives an RTO of minutes and an RPO of seconds.
Data residency
Residency rules turn regions into hard boundaries for some data. Common design: user records and content for EU users live only in EU regions, with only non-personal or aggregated data replicated globally. This interacts with everything else: search indexes, caches, logs, backups and analytics pipelines must respect the same boundary, which is where most violations happen.
Costs people forget
- Cross-region data transfer is billed, and replication traffic adds up fast.
- Idle standby capacity, or headroom in every active region.
- Every operational process (deploys, migrations, incident response) now runs in several places.
In the interview
When asked about region failure, give a proportionate answer: "We run multi-zone in one region by default. For regional survival, the stateless tiers run in two regions behind a global load balancer; the order database replicates asynchronously to the second region with an RPO of a few seconds, and we fail over by promoting it. Payment state uses a consistent multi-region store because losing a payment is not acceptable. We rehearse evacuation quarterly."
Checklist
- Which reason drives multi-region: latency, availability or residency.
- The shape: multi-zone, active-passive, home-region active-active, or multi-writer.
- How users are routed, and how fast failover is.
- Synchronous or asynchronous replication per data type, with RPO and RTO.
- How conflicts are avoided or resolved.
- Headroom, rehearsal, and residency boundaries for every copy of the data.