"Designing read-heavy vs write-heavy systems"
How the read/write ratio shapes a design: caching, replicas, denormalisation and CDNs for read-heavy systems; log-structured storage, batching, partitioning, queues and asynchronous processing for write-heavy ones; and the patterns for systems that are heavy on both.
Reading is half of it. See this used in a real interview: walk through Design a URL Shortener →
One of the first numbers to establish in any design is the ratio of reads to writes. A URL shortener might see 100 redirects per link created; a metrics system writes millions of points per second and reads a tiny fraction of them. The ratio decides where you spend complexity: read-heavy systems precompute and cache; write-heavy systems append, batch and partition. Stating the ratio and drawing the right conclusion is a strong early move in an interview.
Find the ratio
From the requirements and estimates:
- Social feeds, product catalogues, URL shorteners, news sites, maps: read-heavy (often 10:1 to 1,000:1).
- Metrics, logs, IoT telemetry, click tracking, location updates, chat ingestion: write-heavy.
- Some are both: messaging (every message written once, read by recipients), ride-sharing (constant location writes plus frequent nearby queries).
Also note the shape: are reads of recent data or any data? Are writes appends or updates? Is there a hot subset? See back-of-envelope estimation.
Read-heavy toolkit
| Technique | Effect |
|---|---|
| Caching (application, distributed, CDN) | most reads never reach the database. See caching |
| Read replicas | spread remaining reads; accept replication lag. See scaling a relational database |
| Denormalisation and precomputation | store data in the shape reads need; compute on write |
| Materialised views and read models | separate read-optimised stores. See event sourcing and CQRS |
| Fan-out on write | build each reader's view at write time. See fan-out on write vs read |
| Indexes and search engines | fast lookup patterns. See database indexing |
| Edge delivery | static and cacheable content close to users. See CDN and edge |
The trade-off: every copy must be kept fresh, so read-heavy designs move work and complexity to the write path (invalidation, fan-out, index updates). See cache invalidation.
Write-heavy toolkit
| Technique | Effect |
|---|---|
| Append-only, log-structured storage | sequential writes; LSM-based databases. See storage engines |
| Partitioning by a high-cardinality key | spread writes evenly. See sharding and partitioning |
| Batching and buffering | many small writes become few large ones |
| Queues and logs in front of storage | absorb bursts; process at a steady rate. See how Kafka works |
| Pre-aggregation in streams | store per-minute counts instead of every event. See counting at scale |
| Asynchronous secondary work | indexes and derived data updated later |
| Fewer indexes on hot tables | each index multiplies write cost |
| Time-based partitions and TTLs | cheap retention by dropping old data. See time-series data |
The trade-off: reads get harder (data is spread out, not yet aggregated, or eventually consistent), so you add read paths deliberately: rollups for dashboards, a search index fed asynchronously.
Watch the hot spots
- Read-hot keys (a celebrity profile, a viral link): replicate and cache locally.
- Write-hot keys (a global counter, a popular item): shard the counter, batch, or queue per key.
See hot keys and skew.
Heavy on both
Many real systems need both:
- Chat (Design WhatsApp): append messages to a partitioned store keyed by conversation (write path), read recent messages per conversation by range scan (read path), cache recent conversations.
- Ride-sharing: location writes go to an in-memory geospatial index (not a database), with history persisted asynchronously; nearby queries read the in-memory index.
- Feeds: posts are written once, fanned out to timelines (write amplification), read from precomputed timelines (cheap reads).
The usual pattern is CQRS-style separation: an optimised write path that records facts quickly, and one or more read models updated asynchronously for each query shape.
Consistency implications
Read-heavy designs with caches and replicas serve slightly stale data; write-heavy designs with buffering and asynchronous processing show updates later. Decide which reads must be fresh (a user's own writes, balances) and route those to the source of truth. See consistency models.
In the interview
Say the ratio and its consequence early: "Redirects outnumber creations about 100 to 1, so the design centres on caching and edge delivery; writes are small and go to a single partitioned table" (Design a URL Shortener). Or: "We ingest a million points per second and query a tiny fraction, so we append to a log, pre-aggregate in streams, and store in a time-partitioned columnar store" (Design a Monitoring System).
Checklist
- Read/write ratio and access shape estimated first.
- Read-heavy: caches, replicas, CDNs, precomputation, denormalisation.
- Write-heavy: append-only storage, partitioning, batching, queues, pre-aggregation.
- Hot keys handled on whichever side is hot.
- Separate write path and read models when heavy on both.
- Fresh reads routed to the source of truth where needed.