SysDesignPrep.com
Study guide 130 of 183

"Privacy in system design: deletion, retention and residency"

Designing for privacy requirements: knowing where personal data lives, deleting a user everywhere including caches, logs, backups and analytics, retention limits, data residency, minimisation, and consent.

Reading is half of it. See this used in a real interview: walk through Design WhatsApp →

"Delete my account" sounds like one DELETE statement. In a real system that user’s data lives in a dozen databases, caches, search indexes, event streams, logs, analytics tables, backups and third-party services. Laws such as the GDPR and CCPA give users the right to have it removed within weeks, and interviewers increasingly ask how a design handles it. Privacy is far cheaper to design in than to retrofit.

Know where personal data lives

You cannot delete what you cannot find. Keep a data inventory: for each store, which personal data it holds, keyed by what, why, for how long, and who owns it. In a design interview, a sentence like "user id is the key for all personal data, and every store that holds it is registered for deletion" goes a long way.

Design choices that make this easier:

  • Key personal data by user id everywhere, so it can be found.
  • Keep personal fields in a few clearly owned services instead of copying them into every table.
  • Prefer references (user id) in events and logs over copies of names, emails and addresses.

Deleting a user everywhere

A deletion is a workflow, not a query:

  1. Mark the account as deleted immediately (it disappears from the product and stops processing).
  2. Publish a deletion event; every service that holds personal data subscribes and deletes or anonymises its copy, then acknowledges.
  3. Track completion per service, with retries and alerts if a service has not confirmed within the deadline.
  4. Delete from search indexes, caches and CDNs, which are easy to forget.
  5. Handle third parties (email provider, analytics vendor, payment processor) through their deletion APIs.
  6. Record that the deletion happened (without keeping the data) for audit.

Some data must legally be kept (financial records, invoices for tax law, fraud evidence). Separate it, restrict access, and delete it when its own retention period ends.

Logs, analytics and backups

  • Logs: avoid logging personal data in the first place (log ids, not emails); keep logs for short retention periods so deletion happens by expiry.
  • Analytics: pseudonymise identifiers, aggregate early, and make the warehouse’s user tables deletable. See OLTP versus OLAP.
  • Backups: you usually cannot edit them. Common approaches: short backup retention so data ages out, re-applying deletions after any restore, or crypto-shredding (encrypt each user’s data with their own key and delete the key). See encryption and key management.
  • Event logs and append-only stores: compaction with tombstones, crypto-shredding, or keeping personal fields out of the log and referencing them instead. See event sourcing.

Retention

Keep data only as long as it has a purpose: raw location history for 30 days, chat media until delivered or 30 days, logs for 14 days. Enforce retention with TTLs, lifecycle rules and partition drops, not with a cleanup job someone has to remember. Time-partitioned tables make retention cheap. See time-series data.

Minimisation and pseudonymisation

  • Collect only what the feature needs; precise location for a ride, but not stored forever at full precision. See Design Uber.
  • Pseudonymise (replace identifiers with tokens) for analytics and machine learning; keep the mapping separately and access-controlled.
  • Aggregate where possible: counts per region rather than per user.
  • Coarsen data with age: exact location becomes city, exact timestamp becomes day.

Data residency

Some laws and enterprise contracts require data about certain users to stay in a region (the EU, India). Residency affects the whole design:

  • Route those users to in-region infrastructure and store their data there.
  • Make sure every copy stays in region too: replicas, backups, search indexes, caches, logs and support tools.
  • Replicate only non-personal or aggregated data globally.

See multi-region architecture.

Record what each user has consented to (marketing, personalisation, analytics cookies) with a timestamp and version, check it at the point of use, and propagate changes quickly. Use data only for the purpose it was collected for; a new purpose may need new consent.

Access requests

Users can also ask for a copy of their data. The same inventory and events that drive deletion can drive export: each service contributes its part to a downloadable archive within the deadline.

In the interview

When the design holds personal data (location, messages, health, payments), add one sentence on privacy: what is kept and for how long, how a deletion propagates through stores, caches, logs and backups, and whether residency applies. It signals maturity at very little time cost.

Checklist

  • A data inventory keyed by user id.
  • Deletion as an event-driven workflow with tracked completion.
  • Logs and analytics minimised and pseudonymised; backups handled by expiry or crypto-shredding.
  • Retention enforced by TTLs and lifecycle rules.
  • Residency applied to every copy, including backups and indexes.
  • Consent recorded and checked at use.

Open in your browser to sign in

Google does not allow sign-in inside this app's built-in browser. Open this page in Safari and sign in there. The link opens this same page.

Tap the ⋯ or share button at the top or bottom of the screen, then Open in browser. Or copy the link and paste it into Safari.