SysDesignPrep.com
Study guide 07 of 183

Back-of-envelope estimation, worked examples

Five complete capacity estimates for classic interview questions: a URL shortener, a chat app, a video platform, a news feed and a metrics system, with traffic, storage, bandwidth and server counts, and what each number means for the design.

Reading is half of it. See this used in a real interview: walk through Design a URL Shortener →

Estimation is easiest to learn from examples. Below are five quick estimates in the style you would do at a whiteboard: round numbers, a few lines each, and, most importantly, the design conclusion each number leads to. Use them as templates. All numbers are deliberately approximate; the method matters more than the exact assumptions.

Handy conversions: one day is about 86,400 seconds, call it 100,000 (10^5). So 1 million per day is about 12 per second; 1 billion per day is about 12,000 per second. Peak is often 2 to 3 times average. See back-of-envelope estimation.

1. URL shortener

Assumptions: 100 million new links per month; 100 redirects per link created.

  • Writes: 100 M / month ≈ 3.3 M / day ≈ 40 per second (peak about 100).
  • Reads: 100x ≈ 4,000 per second (peak about 10,000).
  • Storage: 500 bytes per link (URL plus metadata) × 100 M per month × 12 × 5 years ≈ 3 TB.
  • Codes: 6 billion links over 5 years; base62 with 7 characters gives 3.5 trillion, plenty.

Conclusions: writes are trivial for one database; reads are cacheable and highly skewed (hot links), so cache and edge-cache redirects; 3 TB fits on a single well-provisioned database, but sharding by code is easy if needed. See Design a URL Shortener.

2. Chat app

Assumptions: 500 million daily active users; 40 messages sent per user per day; average message 100 bytes; 30 % of users online at peak.

  • Messages: 500 M × 40 = 20 billion per day ≈ 230,000 per second average, about 500,000 at peak.
  • Storage: 20 B × 100 bytes = 2 TB per day of text (more with metadata and indexes, say 5 TB per day), about 2 PB per year before replication.
  • Connections: 30 % of 500 M = 150 million concurrent connections; at about 500,000 per gateway server, 300 gateways (plus headroom).
  • Media dominates bytes: even 5 % of messages with a 200 KB photo is 200 TB per day, stored in object storage with a CDN.

Conclusions: a write-heavy, partitioned message store (by conversation), a large fleet of connection gateways, and media separated from messages. See Design WhatsApp.

3. Video platform

Assumptions: 500 hours of video uploaded per minute; 1 billion hours watched per day; average streaming bitrate 3 Mbps.

  • Uploads: 500 × 60 × 24 ≈ 720,000 hours per day. At about 1 GB per hour for the original plus about 2 GB per hour for all renditions, roughly 2 PB per day of new storage.
  • Watch traffic: 1 B hours × 3,600 s × 3 Mbps ≈ 1.1 × 10^19 bits per day ≈ 1.3 exabytes per day; spread over a day, an average egress of about 125 Tbps.
  • Transcoding: if transcoding takes about 2 CPU-hours per video hour per rendition set, 720,000 × 2 ≈ 1.4 M CPU-hours per day ≈ 60,000 cores busy continuously.

Conclusions: delivery must be almost entirely from CDNs and ISP-embedded caches; storage needs tiering (most videos are rarely watched); transcoding runs on a large, preemptible batch fleet. See Design YouTube and cost-aware system design.

4. News feed

Assumptions: 300 million daily active users; each opens the feed 10 times a day; each user posts 0.5 times a day; average 200 followers; feed page of 20 posts.

  • Feed reads: 300 M × 10 = 3 B per day ≈ 35,000 per second, peak about 100,000.
  • Posts: 150 M per day ≈ 1,700 per second.
  • Fan-out on write: 1,700 × 200 = 340,000 timeline inserts per second on average.
  • Timeline cache: 300 M users × 500 post ids × 8 bytes ≈ 1.2 TB of ids in memory, feasible in a Redis cluster.
  • A celebrity with 50 M followers would cost 50 M inserts per post: handle with fan-out on read for such accounts.

Conclusions: precomputed timelines in an in-memory store, hybrid fan-out, posts hydrated from a cache. See Design a News Feed and fan-out on write vs read.

5. Metrics system

Assumptions: 100,000 hosts; 1,000 time series per host; one sample every 10 seconds.

  • Series: 100,000 × 1,000 = 100 million active series.
  • Ingest: 100 M / 10 s = 10 million samples per second.
  • Storage: raw 16 bytes per sample would be 160 MB/s, about 14 TB per day; with time-series compression (about 1.5 bytes per sample) about 1.3 TB per day.
  • Retention: raw for 15 days (about 20 TB), then downsampled to 1-minute and 1-hour resolution for a year at a fraction of the size.

Conclusions: a write-heavy, partitioned ingestion path (by series), specialised compression, and downsampling with tiered retention; queries hit recent data mostly, so keep it in memory or on fast disks. See Design a Monitoring System and time-series data.

How to present an estimate

  1. State assumptions out loud and round aggressively.
  2. Compute QPS (average and peak), storage per year, and bandwidth if media is involved.
  3. Find the largest single unit: the hottest key, the biggest user, the biggest partition.
  4. Say what the numbers imply: "fits on one machine", "needs sharding", "must come from a CDN", "fan-out is the bottleneck".
  5. Move on; two to four minutes is enough.

See latency numbers for the per-operation costs behind server counts.

Checklist

  • One day is about 10^5 seconds; peak is 2 to 3 times average.
  • QPS, storage, bandwidth and concurrency estimated.
  • The largest single key or user identified.
  • Every number tied to a design decision.
  • Media and replication accounted for separately.

Open in your browser to sign in

Google does not allow sign-in inside this app's built-in browser. Open this page in Safari and sign in there. The link opens this same page.

Tap the ⋯ or share button at the top or bottom of the screen, then Open in browser. Or copy the link and paste it into Safari.