Latency numbers every engineer should know
The approximate costs of memory access, disk and SSD reads, network round trips within a data centre and across continents, serialization, database queries and cache hits, with what each implies for system design and how to use them in back-of-envelope estimates.
Reading is half of it. See this used in a real interview: walk through Design a URL Shortener →
Good design decisions depend on knowing roughly how long things take. Is a cache worth it? Can we afford a cross-region call in the request path? How many disk reads fit in a 100 ms budget? The famous "latency numbers every programmer should know" list answers these by orders of magnitude. Exact values vary by hardware and change over time; what matters is the ratios. Here is a modernised version with the conclusions that matter for system design.
The numbers (approximate)
| Operation | Time | Relative |
|---|---|---|
| CPU L1 cache reference | about 1 ns | 1 |
| Branch mispredict | about 3 to 5 ns | |
| L2 cache reference | about 4 ns | |
| Mutex lock or unlock (uncontended) | about 20 ns | |
| Main memory reference | about 100 ns | 100x L1 |
| Compress 1 KB with a fast compressor | about 1 to 2 µs | |
| Read 1 MB sequentially from memory | a few µs to 10 µs | |
| Send 1 KB over a 10 Gbps network | about 1 µs | |
| Random read from NVMe SSD | about 20 to 100 µs | |
| Round trip within a data centre (same zone) | about 100 to 500 µs | |
| Read 1 MB sequentially from NVMe SSD | about 100 to 300 µs | |
| Redis or Memcached GET over the network | about 200 µs to 1 ms | |
| Round trip between zones in one region | about 0.5 to 2 ms | |
| Simple indexed database query (warm cache) | about 1 to 5 ms | |
| Hard disk seek | about 5 to 10 ms | |
| Read 1 MB sequentially from a hard disk | about 5 ms | |
| Round trip across a continent (US east to west) | about 60 to 80 ms | |
| Round trip across an ocean (US to Europe) | about 80 to 150 ms | |
| Round trip US to Australia or Asia | about 150 to 250 ms | |
| TLS handshake to a distant server | 1 to 2 extra round trips |
Lessons for design
Memory is about 1,000 times faster than SSD random reads, which are about 100 times faster than disk seeks. That is why caches and in-memory data structures dominate latency-sensitive designs, and why LSM trees and append-only logs favour sequential I/O. See storage engines.
Within a data centre, the network is fast; across the world, it is slow. A same-zone call costs well under a millisecond, so a few internal service calls are affordable. A cross-ocean round trip costs about 100 ms, and the speed of light cannot be optimised away. Anything in the user's request path should avoid cross-region calls; that is why we use CDNs, regional deployments and edge termination. See multi-region architecture and CDN and edge.
Round trips matter more than bandwidth for small requests. A new HTTPS connection to a distant server costs several round trips before any data flows. Connection reuse, HTTP/2 and HTTP/3, and terminating TLS near the user save hundreds of milliseconds. See what happens when you type a URL.
Consensus and synchronous replication cost round trips. A strongly consistent write across three regions waits for at least one cross-region round trip. Within one region it is a millisecond or two. See Spanner and distributed SQL.
Sequential beats random. Reading 1 MB sequentially from SSD takes about the same time as a handful of random reads; batch and stream data rather than fetching it piecemeal.
Using them in estimates
Example: a feed page must render in 200 ms for users in Europe, served from a European region.
- Network to the region and back: about 20 to 40 ms.
- Budget left for the backend: about 150 ms.
- A cache hit costs about 1 ms; a database query about 5 ms; a call to another service about 1 to 2 ms plus its own work.
- Ten sequential database queries (50 ms) fit, but are risky for the tail; parallelise and cache instead. See tail latency.
- A synchronous call to a US region (about 100 ms round trip) would consume most of the budget: avoid it.
Throughput from latency: one thread doing 5 ms queries sequentially manages about 200 per second; to handle 10,000 per second you need about 50 concurrent queries in flight. See queueing theory and capacity planning.
Throughput numbers worth knowing
| Resource | Rough capacity |
|---|---|
| A single Redis instance | 100,000+ simple operations per second |
| A well-tuned relational database primary | thousands to tens of thousands of transactions per second |
| A Kafka broker | hundreds of MB per second |
| A web server instance | thousands of simple requests per second |
| A 10 Gbps network link | about 1.25 GB per second |
| NVMe SSD | hundreds of thousands of random reads per second; several GB per second sequential |
These vary widely with workload, but they let you sanity-check a design in seconds. See back-of-envelope estimation.
Where this shows up in questions
- Design Typeahead: a keystroke budget of tens of milliseconds forces in-memory lookups and edge caching.
- Design a Stock Exchange: microsecond budgets mean no network hops or disk seeks on the matching path.
- Design a Distributed Cache: sub-millisecond network gets versus milliseconds from the database is the reason it exists.
- Design a URL Shortener: redirects cached at the edge avoid an ocean round trip.
Checklist
- Memory, SSD, disk and network latencies by order of magnitude.
- No cross-region calls in user-facing request paths.
- Minimise round trips: connection reuse, edge termination, batching.
- Prefer sequential I/O and in-memory hot data.
- Convert latency into concurrency with Little's law.
- Sanity-check component throughput with rough capacity numbers.