SysDesignPrep.com
Study guide 180 of 183

Backups and disaster recovery

RPO and RTO, why replication is not a backup, full, incremental and point-in-time backups, testing restores, ransomware and accidental deletes, disaster recovery tiers, and how to answer "what if we lose the database?"

Reading is half of it. See this used in a real interview: walk through Design a Payment System →

"What happens if someone runs a bad migration and deletes half the orders table?" Replication does not help: the delete is faithfully replicated everywhere within milliseconds. Backups and disaster recovery cover the failures that redundancy cannot: human error, software bugs that corrupt data, ransomware, and losing a whole region. They rarely get much time in an interview, which is exactly why a crisp answer stands out.

Two numbers define everything

  • RPO (recovery point objective): how much recent data you can afford to lose. "No more than 5 minutes of orders."
  • RTO (recovery time objective): how long the service can be down while you recover. "Back within 1 hour."

Every choice below trades cost against these two numbers. Ask for them, or propose them per data type: payments might need an RPO near zero; analytics can lose a day.

Replication is not a backup

Replicas protect against hardware failure: a disk dies, a machine dies, a zone dies. They do not protect against:

  • A bad DELETE or a buggy deploy that writes garbage (replicated instantly).
  • Corruption that propagates.
  • An attacker or ransomware with access to the primary.
  • Deleting the wrong database or bucket.

For those you need copies from the past that the live system cannot modify.

Backup types

TypeWhat it isRestore
Full snapshota complete copy at one moment (disk snapshot or logical dump)fast to restore that moment
Incrementalonly what changed since the last backupneeds the chain of increments
Continuous log archivingship the database’s write-ahead log or binlog continuouslyreplay to any second: point-in-time recovery (PITR)

The standard database setup: a daily snapshot plus continuous WAL archiving, which gives point-in-time recovery to any moment in the retention window with an RPO of seconds to minutes. Managed databases (RDS, Cloud SQL, Aurora) provide exactly this.

For object storage: versioning (deletes and overwrites keep the old version) plus lifecycle rules to expire old versions, and replication to another region or account.

Make backups hard to destroy

  • Store backups in a separate account or project with separate credentials, so whoever compromises production cannot delete them.
  • Use immutable storage (object lock, write-once retention) for a retention window.
  • Encrypt them, and keep the keys recoverable separately.
  • Keep at least one copy in another region.

The old rule of thumb still works: 3 copies, 2 different media or services, 1 off-site (and ideally 1 immutable).

A backup you have not restored is a hope

Most backup failures are discovered during the restore: missing files, expired keys, a format nobody can read, a restore that takes 30 hours instead of the 1-hour RTO.

  • Restore regularly and automatically: spin up a database from last night’s backup, run integrity checks and row counts, and alert if it fails.
  • Measure restore time: a 10 TB database restores at a few hundred MB/s at best, which is many hours. If that breaks the RTO, you need a different strategy (a delayed replica, smaller shards, or warm standby).
  • Practise partial restores: recovering one customer’s data or one table without rolling back everyone.

Recovering from logical mistakes

For "someone deleted half the orders table at 14:03":

  1. Stop the damage (revoke access, disable the job).
  2. Restore a copy of the database to a point just before 14:03, beside production, not over it.
  3. Extract the missing or corrupted rows and repair production, preserving writes made since.

A delayed replica (a follower that applies changes an hour late) offers a faster path: stop it before it applies the bad change and copy data from it. Soft deletes and audit logs also make many mistakes reversible without touching backups.

Disaster recovery tiers

StrategyRTORPOCost
Backup and restore in another regionhours to a dayminutes to hourslowest
Pilot light (data replicated, minimal infrastructure running)tens of minutes to hoursseconds to minuteslow
Warm standby (scaled-down full copy running)minutessecondsmedium
Active-active multi-regionnear zeronear zerohighest

Choose per service: the checkout path may need warm standby; the internal reporting tool can live with backup and restore. See multi-region architecture.

Do not forget

  • Configuration and infrastructure (infrastructure as code, secrets, DNS) are part of recovery; a restored database is useless without them.
  • Dependencies: if your DR region relies on a service only available in the primary region, it is not DR.
  • Runbooks and drills: a written, rehearsed procedure turns a 6-hour panic into a 45-minute recovery.
  • Retention and privacy: backups contain personal data; deletion requests and retention limits apply to them too.

In the interview

"We take daily snapshots with continuous WAL archiving for point-in-time recovery, kept for 30 days in a separate account with object lock. Restores are tested nightly. For a bad migration we restore beside production and repair. For regional loss, the order database has a warm standby in a second region with an RPO of seconds and an RTO of about 15 minutes."

Checklist

  • RPO and RTO per data type.
  • Snapshots plus log archiving for point-in-time recovery.
  • Backups isolated from production credentials, immutable, off-site.
  • Automated restore tests and measured restore time.
  • A procedure for logical errors that does not overwrite newer data.
  • A DR tier per service, with configuration and dependencies included.

Open in your browser to sign in

Google does not allow sign-in inside this app's built-in browser. Open this page in Safari and sign in there. The link opens this same page.

Tap the ⋯ or share button at the top or bottom of the screen, then Open in browser. Or copy the link and paste it into Safari.