System design patterns cheat sheet
A one-page map of the patterns that solve recurring system design problems: scaling reads and writes, consistency, reliability, messaging, data processing, real-time, search and geo, security and operations, each with a one-line use and a link to the full guide.
Reading is half of it. See this used in a real interview: walk through Design a URL Shortener →
Most system design questions are combinations of a few dozen recurring problems, each with well-known solutions. This page lists them by problem, with a one-line reminder of when to use each pattern and a link to the guide that explains it properly. Use it to review before an interview, or to find what to read next.
Scaling reads
| Problem | Pattern | Guide |
|---|---|---|
| Same data read over and over | cache-aside with TTLs | caching |
| Static or public content worldwide | CDN with long-lived, versioned URLs | CDN and edge |
| Database read load | read replicas, lag-aware routing | scaling a relational database |
| Expensive queries per page | denormalise, precompute read models | data modelling and denormalisation |
| Feeds and timelines | fan-out on write, pull for celebrities | fan-out on write vs read |
| One extremely hot key | local cache, replicated keys, request coalescing | hot keys and skew |
| Stale cache after updates | delete on write, CDC invalidation, leases | cache invalidation |
Scaling writes and data
| Problem | Pattern | Guide |
|---|---|---|
| Data or writes beyond one machine | shard by a high-cardinality key | sharding and partitioning |
| Adding nodes without reshuffling everything | consistent hashing | consistent hashing |
| Very high write volume | LSM storage, batching, append-only logs | storage engines |
| Counting likes, views, clicks | sharded counters, stream aggregation | counting at scale |
| Metrics and events over time | time partitions, downsampling, TTLs | time-series data |
| Large files and media | object storage, presigned uploads | media uploads and processing |
| Unique ids across machines | Snowflake-style or UUIDv7 ids | unique ids, ordering and time |
Consistency and correctness
| Problem | Pattern | Guide |
|---|---|---|
| Choosing guarantees per feature | linearizable, causal, read-your-writes, eventual | consistency models |
| Concurrent updates to one record | conditional updates, optimistic versions, locks | optimistic vs pessimistic locking |
| No double booking or overselling | constraints, holds with expiry | bookings and reservations |
| Retries causing duplicates | idempotency keys | distributed transactions and idempotency |
| Multi-service transactions | sagas with compensation | workflow orchestration |
| Money movement | double-entry ledger, reconciliation | payments and ledgers |
| One leader or lock holder | leases with fencing tokens, consensus | distributed locks and leases |
| Agreement across replicas | Raft or Paxos | how Raft works |
Messaging and asynchrony
| Problem | Pattern | Guide |
|---|---|---|
| Slow side effects in the request path | queue them, respond early | sync vs async communication |
| Many consumers of the same events | a partitioned log (Kafka) | how Kafka works |
| Database change and event must both happen | transactional outbox, CDC | change data capture |
| At-least-once delivery | idempotent consumers, dead-letter queues | delivery semantics |
| Work at a future time | durable timers, partitioned schedulers | delayed jobs and distributed cron |
Reliability
| Problem | Pattern | Guide |
|---|---|---|
| Dependency failures cascading | timeouts, retries with jitter, circuit breakers | rate limiting and resilience |
| Traffic spikes beyond capacity | load shedding, backpressure, waiting rooms | load shedding and backpressure |
| Abuse and overuse | token bucket or sliding window limits | rate limiting algorithms |
| Losing a zone or region | active-active services, standby databases | active-active vs active-passive |
| Global users | multi-region with home regions | multi-region architecture |
| Data loss or corruption | backups, point-in-time restore, tested recovery | backups and disaster recovery |
| Slow outliers | hedged requests, parallelism, deadlines | tail latency |
Real-time and collaboration
| Problem | Pattern | Guide |
|---|---|---|
| Pushing updates to clients | WebSockets, SSE or polling | WebSockets vs SSE vs long polling |
| Millions of connections and presence | gateways, session registry, TTL presence | presence and connection management |
| Reaching users when the app is closed | push notifications | push notifications |
| Concurrent editing | CRDTs or operational transformation | CRDTs vs operational transformation |
| Offline use and sync | local-first storage, change logs, conflict rules | offline-first apps and sync |
Search, ranking and geo
| Problem | Pattern | Guide |
|---|---|---|
| Full-text search | inverted index, BM25 | how Elasticsearch works |
| Ordering results well | retrieve then rank, learning to rank | search ranking and relevance |
| Search as you type | top-K per prefix | tries and autocomplete |
| Nearby things | geohash, S2, H3 or quadtrees | geohash vs quadtree vs H3 |
| Matching supply and demand | dispatch with ETA and atomic claims | dispatch and marketplace matching |
| Personalised recommendations | candidate generation plus ranking | recommendation systems |
| Meaning-based retrieval | embeddings and vector indexes | vector search and RAG |
Data processing
| Problem | Pattern | Guide |
|---|---|---|
| Aggregating streams correctly | event-time windows, watermarks | windowing and watermarks |
| Large batch computations | MapReduce or Spark | MapReduce and Spark |
| Analytical storage | columnar lakehouse tables | data lakes and lakehouses |
| Approximate answers at scale | Bloom filters, HyperLogLog, count-min | probabilistic data structures |
Security and operations
| Problem | Pattern | Guide |
|---|---|---|
| Login and sessions | short tokens, refresh rotation | JWT vs session cookies |
| Who can access what | RBAC plus relationship-based permissions | authorization and permissions |
| Many customers on shared infrastructure | tenant ids, quotas, cells | multi-tenancy |
| Knowing the system is healthy | SLOs and burn-rate alerts | SLIs, SLOs and error budgets |
| Debugging across services | distributed tracing | distributed tracing |
| Shipping safely | canaries, flags, progressive rollout | deployment strategies |
Putting it together
In an interview, identify the two or three problems that make the question hard, then reach for the matching patterns and explain their trade-offs. For example, Design a News Feed is mostly fan-out, caching and ranking; Design WhatsApp is connections, delivery guarantees and ordering; Design a Payment System is idempotency, ledgers and sagas; Design a URL Shortener is id generation, caching and edge delivery. Start with the interview framework for the overall structure.
Checklist
- Name the hard problems in the question first.
- Pick the matching pattern for each, with its trade-off.
- Keep the rest of the design simple.
- Read the linked guide for any pattern you cannot explain in two sentences.