SysDesignPrep.com
Study guide 84 of 183

"Active-active vs active-passive"

High-availability topologies compared: active-passive with hot, warm and cold standbys, active-active within and across regions, failover detection and promotion, split brain and fencing, write conflicts, RTO and RPO, cost, and how to choose for databases, services and whole regions.

Reading is half of it. See this used in a real interview: walk through Design a Payment System →

To survive failures, you run more than one copy of everything. The question is whether the copies all serve traffic at once (active-active) or one serves while the others wait (active-passive). The choice shapes failover speed, data consistency, cost and complexity, and it applies at every level: a database, a service, or an entire region. Interviewers ask "what happens when this fails?"; your answer usually names one of these topologies.

Active-passive

One primary handles all traffic; one or more standbys receive replicated data and take over if the primary fails.

Standby temperature:

  • Hot standby: running, fully replicated, ready to take over in seconds to a minute.
  • Warm standby: running at reduced capacity, needs scaling up before taking full load; minutes.
  • Cold standby: infrastructure defined but not running; restore from backups; hours.

Properties:

  • Simple consistency: one writer, so no write conflicts.
  • Wasted capacity: standbys mostly sit idle (though hot database standbys can often serve reads).
  • Failover is an event: detection, promotion, and redirecting traffic take time, and failover paths that are rarely exercised tend to break when needed.

Typical: relational database primary with a synchronous or asynchronous replica in another zone; a disaster-recovery region for less critical systems.

Active-active

All copies serve traffic simultaneously; load balancers or DNS spread users across them.

Properties:

  • No idle capacity, and failover is just "stop sending traffic to the failed copy", usually fast and well tested because the path is used constantly.
  • Each copy must have headroom to absorb the others' traffic when one fails.
  • For stateless services, active-active is the default and easy.
  • For data, it is hard: if two copies accept writes to the same record, they can conflict.

Writes in active-active

Ways to handle writes when several sites are active:

ApproachHowTrade-off
Partition ownership (home region)each user or record is written in one region only; others forward or read replicasno conflicts; cross-region writes for travelling users
Consensus across siteswrites commit on a majority of regionsstrong consistency; cross-region latency per write
Multi-leader with conflict resolutionany region accepts writes; replicate asynchronously and resolve conflictslow latency; last-writer-wins loses data, merges need design
CRDTsdata types that merge automaticallyonly for data that fits those types

See multi-region architecture, Spanner and distributed SQL and CRDTs vs operational transformation.

Failover mechanics

  • Detection: health checks and failure detectors; require several failures to avoid flapping. See gossip and failure detection.
  • Promotion: choose the most up-to-date standby; with consensus-based systems this is automatic.
  • Fencing: ensure the old primary cannot keep accepting writes (revoke its lease, block it at the network, use epoch numbers), or you get split brain: two primaries diverging. See distributed locks and leases.
  • Redirect traffic: update the endpoint, DNS or load balancer; clients reconnect. See global traffic management.

RTO and RPO

  • RTO (recovery time objective): how long until service is back.
  • RPO (recovery point objective): how much recent data you can lose.
TopologyTypical RTOTypical RPO
Cold standby from backupshourssince the last backup
Warm standby, async replicationminutesseconds of writes
Hot standby, async replicationunder a minuteseconds
Hot standby, synchronous replicationunder a minutezero
Active-active with consensussecondszero

Synchronous replication gives zero data loss at the cost of write latency, which is why it is usually done within a region (across zones) rather than across continents. See backups and disaster recovery.

Cost

Active-passive across regions doubles infrastructure for little daily use; active-active uses everything but requires each site to have spare capacity (with three active regions, each must handle half the total load if one fails) and much more engineering for data. See cost-aware system design.

Choosing

  • Stateless services: active-active across zones always; across regions when latency or availability goals justify it.
  • Relational databases: active-passive with a hot synchronous standby in another zone; asynchronous replica in another region for disaster recovery.
  • Global, read-heavy systems: active-active reads everywhere, writes to a home region. See Design a URL Shortener.
  • Money and inventory: single-writer per record (home region or consensus), never last-writer-wins. See Design a Payment System.
  • Chat and social: users homed to a region, active-active overall. See Design WhatsApp.

In the interview

"Application servers are active-active across three zones. The database is active-passive with a synchronous standby in another zone, giving sub-minute failover and no data loss; fencing prevents split brain. A warm standby region with asynchronous replication covers regional disasters with an RPO of seconds." For a key-value store, explain how replicas share writes (quorums or consensus) instead.

Checklist

  • Topology chosen per component: services, databases, regions.
  • Standby temperature matched to RTO.
  • Write strategy for active-active data: ownership, consensus, or resolved conflicts.
  • Detection, promotion, fencing and traffic redirection planned and practised.
  • RTO and RPO stated; synchronous replication where RPO must be zero.
  • Capacity headroom to absorb a failed site.

Open in your browser to sign in

Google does not allow sign-in inside this app's built-in browser. Open this page in Safari and sign in there. The link opens this same page.

Tap the ⋯ or share button at the top or bottom of the screen, then Open in browser. Or copy the link and paste it into Safari.