Backups and disaster recovery
RPO and RTO, why replication is not a backup, full, incremental and point-in-time backups, testing restores, ransomware and accidental deletes, disaster recovery tiers, and how to answer "what if we lose the database?"
Reading is half of it. See this used in a real interview: walk through Design a Payment System →
"What happens if someone runs a bad migration and deletes half the orders table?" Replication does not help: the delete is faithfully replicated everywhere within milliseconds. Backups and disaster recovery cover the failures that redundancy cannot: human error, software bugs that corrupt data, ransomware, and losing a whole region. They rarely get much time in an interview, which is exactly why a crisp answer stands out.
Two numbers define everything
- RPO (recovery point objective): how much recent data you can afford to lose. "No more than 5 minutes of orders."
- RTO (recovery time objective): how long the service can be down while you recover. "Back within 1 hour."
Every choice below trades cost against these two numbers. Ask for them, or propose them per data type: payments might need an RPO near zero; analytics can lose a day.
Replication is not a backup
Replicas protect against hardware failure: a disk dies, a machine dies, a zone dies. They do not protect against:
- A bad
DELETEor a buggy deploy that writes garbage (replicated instantly). - Corruption that propagates.
- An attacker or ransomware with access to the primary.
- Deleting the wrong database or bucket.
For those you need copies from the past that the live system cannot modify.
Backup types
| Type | What it is | Restore |
|---|---|---|
| Full snapshot | a complete copy at one moment (disk snapshot or logical dump) | fast to restore that moment |
| Incremental | only what changed since the last backup | needs the chain of increments |
| Continuous log archiving | ship the database’s write-ahead log or binlog continuously | replay to any second: point-in-time recovery (PITR) |
The standard database setup: a daily snapshot plus continuous WAL archiving, which gives point-in-time recovery to any moment in the retention window with an RPO of seconds to minutes. Managed databases (RDS, Cloud SQL, Aurora) provide exactly this.
For object storage: versioning (deletes and overwrites keep the old version) plus lifecycle rules to expire old versions, and replication to another region or account.
Make backups hard to destroy
- Store backups in a separate account or project with separate credentials, so whoever compromises production cannot delete them.
- Use immutable storage (object lock, write-once retention) for a retention window.
- Encrypt them, and keep the keys recoverable separately.
- Keep at least one copy in another region.
The old rule of thumb still works: 3 copies, 2 different media or services, 1 off-site (and ideally 1 immutable).
A backup you have not restored is a hope
Most backup failures are discovered during the restore: missing files, expired keys, a format nobody can read, a restore that takes 30 hours instead of the 1-hour RTO.
- Restore regularly and automatically: spin up a database from last night’s backup, run integrity checks and row counts, and alert if it fails.
- Measure restore time: a 10 TB database restores at a few hundred MB/s at best, which is many hours. If that breaks the RTO, you need a different strategy (a delayed replica, smaller shards, or warm standby).
- Practise partial restores: recovering one customer’s data or one table without rolling back everyone.
Recovering from logical mistakes
For "someone deleted half the orders table at 14:03":
- Stop the damage (revoke access, disable the job).
- Restore a copy of the database to a point just before 14:03, beside production, not over it.
- Extract the missing or corrupted rows and repair production, preserving writes made since.
A delayed replica (a follower that applies changes an hour late) offers a faster path: stop it before it applies the bad change and copy data from it. Soft deletes and audit logs also make many mistakes reversible without touching backups.
Disaster recovery tiers
| Strategy | RTO | RPO | Cost |
|---|---|---|---|
| Backup and restore in another region | hours to a day | minutes to hours | lowest |
| Pilot light (data replicated, minimal infrastructure running) | tens of minutes to hours | seconds to minutes | low |
| Warm standby (scaled-down full copy running) | minutes | seconds | medium |
| Active-active multi-region | near zero | near zero | highest |
Choose per service: the checkout path may need warm standby; the internal reporting tool can live with backup and restore. See multi-region architecture.
Do not forget
- Configuration and infrastructure (infrastructure as code, secrets, DNS) are part of recovery; a restored database is useless without them.
- Dependencies: if your DR region relies on a service only available in the primary region, it is not DR.
- Runbooks and drills: a written, rehearsed procedure turns a 6-hour panic into a 45-minute recovery.
- Retention and privacy: backups contain personal data; deletion requests and retention limits apply to them too.
In the interview
"We take daily snapshots with continuous WAL archiving for point-in-time recovery, kept for 30 days in a separate account with object lock. Restores are tested nightly. For a bad migration we restore beside production and repair. For regional loss, the order database has a warm standby in a second region with an RPO of seconds and an RTO of about 15 minutes."
Checklist
- RPO and RTO per data type.
- Snapshots plus log archiving for point-in-time recovery.
- Backups isolated from production credentials, immutable, off-site.
- Automated restore tests and measured restore time.
- A procedure for logical errors that does not overwrite newer data.
- A DR tier per service, with configuration and dependencies included.