Back-of-envelope estimator for system design interviews
Enter the givens and get requests per second, storage, bandwidth, cache size and server count, each with the arithmetic written out the way you would say it in the interview. Then read what the numbers imply for the design.
10 M users × 2 a day = 20 M writes a day. ÷ 86 400 s (call it 10^5) = 231/s; × 3 peak = 694/s.
20 M writes a day × 10 reads each = 200 M reads a day = 2.3 k/s; × 3 peak = 6.9 k/s.
20 M writes a day × 1 KB = 20 GB a day; × 365 = 7.3 TB a year.
7.3 TB a year × 5 years × 3 replicas = 110 TB, before indexes (add up to 2× for those).
In: 694/s × 1 KB = 694 KB/s (5.6 Mbit/s). Out: 6.9 k/s × 1 KB = 6.9 MB/s (56 Mbit/s).
200 M reads a day × 1 KB × 20 % hot = 40 GB. An upper bound: repeated reads of the same item are cached once.
(694/s + 6.9 k/s) ÷ 10 k per server = 1, before headroom; run at about half of that capacity.
What it means for the design
- Read-heavy (10 : 1): the read path is the design. Caching and read replicas come before any write sharding.
- 694/s writes at peak fits one well-sized relational primary; writes are not the hard part.
- 110 TB will not fit on a few machines: plan the sharding key now.
- The hot set (40 GB) fits in a small Redis or Memcached cluster.
The method, the numbers worth memorising, and the per-server capacities behind these defaults are in the back-of-envelope estimation guide. Nobody grades the third significant figure: what an interviewer listens for is whether you notice which part of the system the numbers put under pressure.