SysDesignPrep.com
Study guide 83 of 183

Multi-region architecture

Why and when to run in several regions, active-passive versus active-active, routing users to a region, replicating data across regions, conflict handling, failover and evacuation, and data residency rules.

Reading is half of it. See this used in a real interview: walk through Design Netflix →

"What happens if a whole region goes down?" is one of the most common late-interview questions, and "multi-region" is one of the most expensive answers to give without thinking. Running in several regions buys lower latency for distant users and survival of a regional outage, and it costs consistency, money and complexity. This guide is about choosing the right amount of it.

Why go multi-region

  1. Latency: users in Sydney should not pay 200 ms round trips to Virginia for every request.
  2. Availability: survive the loss of an entire region (rare, but it happens, and some customers require it).
  3. Data residency: laws or contracts may require some users’ data to stay in a country or bloc (the EU, India).

Be clear which of these drives the design, because they lead to different architectures. Latency alone is often solved by a CDN and edge caching without multi-region data. See CDNs and edge computing.

The main shapes

Active-passive with one writing region and a standby, compared with active-active where every region serves reads and writes and replicates to the others
Active-passive keeps one writer and a standby; active-active lets every region write and must reconcile.
ShapeWrites go toFailoverConsistencyCost and complexity
Single region, multi-zoneone regionzone failures onlysimple, stronglowest
Active-passive (warm standby)the primary regionpromote the standby; minutesstrong in primary; standby lagsmoderate; standby capacity is idle
Active-active, home region per usereach user’s home regionmove users to another regionstrong per user; cross-user data eventualhigh
Active-active, multi-writerany regioninstant; regions are peerseventual with conflict resolution, or slow global consensushighest

Most companies should start with multi-zone in one region (zones are separate data centres with independent power, a millisecond apart), which already survives the most common failures. Add regions when one of the three reasons demands it.

Routing users to a region

  • Geo DNS or latency-based DNS sends users to the nearest healthy region; slow to change because of DNS caching.
  • Anycast with a global load balancer (Cloudflare, AWS Global Accelerator, Google’s load balancer) routes to the nearest region and fails over in seconds.
  • Home-region routing: the user’s id or tenant maps to a home region, and the edge forwards their writes there, even if they are travelling.

See networking for system design for why DNS alone is not fast failover.

Replicating data across regions

Cross-region round trips are 50 to 150 ms, so the core choice is whether a write waits for other regions.

  • Asynchronous replication: the write commits locally and is shipped to other regions within a second or so. Fast; on regional failure the last moments of writes may be lost (RPO of seconds).
  • Synchronous or consensus replication (Spanner, CockroachDB, a Raft group spanning regions): the write waits for a majority of regions. No data loss on failover; every write pays cross-region latency.

Pick per data type, as the CAP and PACELC guide explains: balances and inventory may justify synchronous replication; feeds, profiles and analytics do not.

Conflicts in active-active

If two regions can accept writes to the same record, they will sometimes do so concurrently. Options, best first:

  1. Avoid conflicts by design: give each record a home region (a user’s data lives in their home region; other regions forward writes there). Most "active-active" systems at large companies are really this.
  2. Merge automatically: CRDTs for counters, sets and text; append-only data that never conflicts.
  3. Last-writer-wins by timestamp: simple, but silently drops concurrent writes; acceptable only where that is harmless.
  4. Surface conflicts to the application or user to resolve.

Failover and evacuation

  • Detect: health checks from several outside vantage points, not just from inside the failing region.
  • Decide: automate failover for stateless tiers; for databases, automatic promotion risks split brain, so many teams require a human decision or use consensus-based stores that handle it.
  • Shift traffic gradually and watch the receiving regions, which must have headroom to absorb the load (running each region at under 60 % or so).
  • Fail back deliberately once the region is healthy and caught up.

The most important practice is rehearsal: evacuate a region on purpose, regularly. Netflix’s region evacuations are the well-known example. A failover path that has never been exercised usually fails when needed.

Know your targets: RTO (how long until service is restored) and RPO (how much recent data may be lost). Asynchronous replication with automated failover typically gives an RTO of minutes and an RPO of seconds.

Data residency

Residency rules turn regions into hard boundaries for some data. Common design: user records and content for EU users live only in EU regions, with only non-personal or aggregated data replicated globally. This interacts with everything else: search indexes, caches, logs, backups and analytics pipelines must respect the same boundary, which is where most violations happen.

Costs people forget

  • Cross-region data transfer is billed, and replication traffic adds up fast.
  • Idle standby capacity, or headroom in every active region.
  • Every operational process (deploys, migrations, incident response) now runs in several places.

In the interview

When asked about region failure, give a proportionate answer: "We run multi-zone in one region by default. For regional survival, the stateless tiers run in two regions behind a global load balancer; the order database replicates asynchronously to the second region with an RPO of a few seconds, and we fail over by promoting it. Payment state uses a consistent multi-region store because losing a payment is not acceptable. We rehearse evacuation quarterly."

Checklist

  • Which reason drives multi-region: latency, availability or residency.
  • The shape: multi-zone, active-passive, home-region active-active, or multi-writer.
  • How users are routed, and how fast failover is.
  • Synchronous or asynchronous replication per data type, with RPO and RTO.
  • How conflicts are avoided or resolved.
  • Headroom, rehearsal, and residency boundaries for every copy of the data.

Open in your browser to sign in

Google does not allow sign-in inside this app's built-in browser. Open this page in Safari and sign in there. The link opens this same page.

Tap the ⋯ or share button at the top or bottom of the screen, then Open in browser. Or copy the link and paste it into Safari.