Cost-aware system design
Designing systems that are cheap to run as well as fast: estimating cost in interviews, where cloud bills come from (compute, storage, egress, managed services), storage tiering, caching and CDNs for cost, right-sizing, spot capacity and unit economics.
Reading is half of it. See this used in a real interview: walk through Design YouTube →
At scale, cost is a design constraint as real as latency. Storing every video in five resolutions, keeping a year of metrics at full resolution or running GPUs around the clock can cost more than the business earns. Senior interviewers increasingly ask "roughly what would this cost?" or "how would you make it cheaper?". You do not need exact prices, but you should know where money goes and which levers move it.
Where the bill comes from
| Category | Typical drivers | Main levers |
|---|---|---|
| Compute | instance hours, GPUs, idle capacity | right-sizing, autoscaling, spot, efficient code |
| Storage | bytes stored, replicas, retention | tiering, compression, retention limits, deduplication |
| Network | egress to the internet, cross-region and cross-zone traffic | CDN, compression, keeping chatty traffic in one zone |
| Managed services | per-request pricing, provisioned throughput | batching, caching, choosing the right service |
| Operations | engineers' time | simplicity, managed services where they save more than they cost |
Egress is the surprise for many teams: serving video or downloads from origin storage directly can cost more than the storage itself. See CDN and edge.
Estimate it
A quick cost estimate builds on your capacity numbers:
- Storage: daily new data multiplied by retention multiplied by replication, then priced per tier.
- Bandwidth: peak and average egress per month.
- Compute: required cores or GPUs at peak, with headroom, and how much of the day they are actually busy.
Even order-of-magnitude numbers ("about 2 PB a year, so tiering matters") change the design conversation. See back-of-envelope estimation.
Storage levers
- Tiering: hot data on SSD or standard object storage, warm on infrequent-access tiers, cold in archive tiers that cost a fraction but take minutes or hours to read. Lifecycle rules move objects automatically by age. See object storage and files.
- Retention: delete what you do not need; downsample old metrics (per-second data becomes per-hour after a week). See time-series data.
- Compression and columnar formats: often 5 to 10 times smaller for logs and analytics.
- Deduplication: content-addressed storage stores identical files and chunks once. See Design Dropbox.
- Erasure coding instead of three full replicas for cold data: similar durability at about 1.5 times the size.
Compute levers
- Right-size: most instances run at low utilisation; match instance size to actual use.
- Autoscale to follow demand, and scale non-urgent work to zero when idle. See autoscaling.
- Spot or preemptible capacity for interruptible work (batch jobs, transcoding, training) at a large discount; design jobs to checkpoint and retry.
- Commitments (reserved or committed-use discounts) for the steady baseline.
- Shift work in time: run batch jobs off-peak on capacity you already pay for.
- Efficiency: a faster serialization format or a better query can remove whole fleets.
Do less work
- Cache expensive results; a cache hit is far cheaper than recomputing or re-querying. See caching.
- Compute on demand instead of upfront when most outputs are never used, such as transcoding rarely watched videos into every format only when requested. See Design YouTube.
- Sample high-volume telemetry such as traces. See Design a Monitoring System.
- Batch small requests into bigger ones to cut per-request charges.
GPUs and AI workloads
GPU time dominates AI system costs. Levers: batching requests to raise utilisation, smaller or quantised models where quality allows, routing easy requests to cheaper models, caching repeated prompts and prefixes, and autoscaling on queue depth. See LLM systems and Design LLM Inference.
Unit economics
Express cost per unit of value: per active user per month, per video hour streamed, per million requests, per order. It shows whether growth makes the business healthier or worse, and which feature is expensive. Tag resources by service and team so costs can be attributed.
Trade-offs to say out loud
- Cheaper storage tiers mean slower and sometimes costly retrieval.
- Spot capacity means interruptions.
- Aggressive caching means staleness.
- Fewer replicas means less durability or availability.
- Engineering time is also cost: a managed service that costs more per request can still be cheaper overall.
Checklist
- A rough monthly cost estimate for storage, bandwidth and compute.
- CDN for egress-heavy content.
- Lifecycle tiering, retention limits, compression and deduplication.
- Right-sizing, autoscaling, spot for interruptible work, commitments for baseline.
- Caching, on-demand computation and sampling to avoid work.
- Cost per unit of value, attributed to teams and features.