Encryption and key management
Encryption in transit, at rest and end to end, envelope encryption with a KMS, key rotation, hashing passwords, tokenisation, crypto-shredding for deletion, and what each one actually protects against.
Reading is half of it. See this used in a real interview: walk through Design WhatsApp →
"We encrypt everything" is a common line in system design answers and almost never enough. Interviewers want to know encrypted where, with which keys, held by whom, and against which attacker. This guide covers the layers and the one idea that makes encryption at scale manageable: envelope encryption.
Three places encryption applies
| Layer | Protects against | Does not protect against |
|---|---|---|
| In transit (TLS) | eavesdropping and tampering on the network | a compromised server, which sees plaintext |
| At rest (disk, database, object storage) | stolen disks, backups or snapshots leaking | anyone who can query the running system |
| End to end | the service operator and anyone who breaches the servers | compromised user devices |
Most systems need the first two everywhere; end-to-end encryption is a product decision with large consequences.
In transit
TLS on every external connection, terminated close to the user (see networking), and TLS between internal services too in a zero-trust network, usually mutual TLS handled by a service mesh so services authenticate each other with certificates. HSTS forces browsers onto HTTPS.
At rest and envelope encryption
Encrypting terabytes directly with a single master key is impractical and dangerous. Envelope encryption solves it:
- Data is encrypted with a data encryption key (DEK), a random symmetric key (AES-256-GCM), often one per object, per file or per tenant.
- The DEK is itself encrypted with a key encryption key (KEK) held in a key management service (AWS KMS, Google Cloud KMS, HashiCorp Vault, or a hardware security module).
- The encrypted DEK is stored next to the data. To read, the service asks the KMS to decrypt the DEK, then decrypts the data locally.
Why it works: bulk encryption happens locally and fast; the master key never leaves the KMS; access to the KMS is logged and controlled by policy; and rotating the master key only re-encrypts small DEKs, not all the data.
Managed databases and object stores do this for you ("encryption at rest with a customer-managed key"). Add application-level encryption for especially sensitive fields (national ids, medical notes) so that even people with database access see ciphertext.
Key rotation and access
- Rotate KEKs on a schedule (yearly or on compromise); keep old versions to decrypt existing data until it is re-wrapped.
- Separate duties: the team that runs the database should not also control the keys that decrypt it.
- Per-tenant keys let a customer revoke access to their data (bring your own key) and limit the blast radius of a leak.
- Audit: every decrypt call to the KMS is logged, which turns key access into evidence.
Passwords are hashed, not encrypted
Passwords must never be decryptable. Store a slow, salted hash (bcrypt, scrypt or Argon2), so a stolen database cannot be reversed quickly even with GPUs. Better still, avoid storing passwords at all (passkeys, an identity provider). See authentication and authorisation.
Tokenisation
For card numbers and similar data, the safest design is not to hold them. A provider (or an internal vault) stores the real value and returns a token; your systems store and pass the token. Only the vault can map it back. This keeps most of your systems out of PCI scope and limits what a breach can expose. See money in system design.
End-to-end encryption
With end-to-end encryption, keys exist only on users’ devices; the server stores and relays ciphertext it cannot read. It changes the system substantially:
- No server-side search, content moderation or history sync from the server.
- Key distribution becomes a core service (published public keys, prekeys), and users need a way to verify keys.
- Multi-device support means encrypting for every device.
- Lost devices can mean lost data unless users keep encrypted backups.
Messaging apps use the Signal protocol for this. See Design WhatsApp. Choose it when the threat model includes the operator itself; otherwise, at-rest plus in-transit encryption with strict access controls is usually the right trade.
Crypto-shredding
Deleting data from every backup and log is hard. If each user’s (or tenant’s) data is encrypted with its own key, deleting the key makes all copies unreadable, including those in backups. It is a practical way to honour deletion requests in immutable or append-only stores. See privacy: deletion, retention and residency.
Common mistakes
- Rolling your own cryptography or modes; use vetted libraries and authenticated encryption (AES-GCM, ChaCha20-Poly1305).
- Keys in source code, config files or environment variables checked into repositories; use a secrets manager.
- Encrypting at rest and then logging plaintext.
- Claiming end-to-end encryption while the server can fetch the keys.
In the interview
Say TLS everywhere (mTLS inside), encryption at rest with envelope encryption and a KMS, application-level encryption or tokenisation for the most sensitive fields, hashed passwords, and who can access the keys. Mention end-to-end encryption only if the product needs it, along with what it costs.
Checklist
- TLS externally and mTLS internally.
- At-rest encryption with envelope encryption and a KMS.
- Field-level encryption or tokenisation for the most sensitive data.
- Key rotation, separation of duties and audit of key use.
- Slow salted password hashes.
- End-to-end encryption only when the threat model demands it.
- Crypto-shredding for deletion.