SysDesignPrep.com
Study guide 08 of 183

Latency numbers every engineer should know

The approximate costs of memory access, disk and SSD reads, network round trips within a data centre and across continents, serialization, database queries and cache hits, with what each implies for system design and how to use them in back-of-envelope estimates.

Reading is half of it. See this used in a real interview: walk through Design a URL Shortener →

Good design decisions depend on knowing roughly how long things take. Is a cache worth it? Can we afford a cross-region call in the request path? How many disk reads fit in a 100 ms budget? The famous "latency numbers every programmer should know" list answers these by orders of magnitude. Exact values vary by hardware and change over time; what matters is the ratios. Here is a modernised version with the conclusions that matter for system design.

The numbers (approximate)

OperationTimeRelative
CPU L1 cache referenceabout 1 ns1
Branch mispredictabout 3 to 5 ns
L2 cache referenceabout 4 ns
Mutex lock or unlock (uncontended)about 20 ns
Main memory referenceabout 100 ns100x L1
Compress 1 KB with a fast compressorabout 1 to 2 µs
Read 1 MB sequentially from memorya few µs to 10 µs
Send 1 KB over a 10 Gbps networkabout 1 µs
Random read from NVMe SSDabout 20 to 100 µs
Round trip within a data centre (same zone)about 100 to 500 µs
Read 1 MB sequentially from NVMe SSDabout 100 to 300 µs
Redis or Memcached GET over the networkabout 200 µs to 1 ms
Round trip between zones in one regionabout 0.5 to 2 ms
Simple indexed database query (warm cache)about 1 to 5 ms
Hard disk seekabout 5 to 10 ms
Read 1 MB sequentially from a hard diskabout 5 ms
Round trip across a continent (US east to west)about 60 to 80 ms
Round trip across an ocean (US to Europe)about 80 to 150 ms
Round trip US to Australia or Asiaabout 150 to 250 ms
TLS handshake to a distant server1 to 2 extra round trips

Lessons for design

Memory is about 1,000 times faster than SSD random reads, which are about 100 times faster than disk seeks. That is why caches and in-memory data structures dominate latency-sensitive designs, and why LSM trees and append-only logs favour sequential I/O. See storage engines.

Within a data centre, the network is fast; across the world, it is slow. A same-zone call costs well under a millisecond, so a few internal service calls are affordable. A cross-ocean round trip costs about 100 ms, and the speed of light cannot be optimised away. Anything in the user's request path should avoid cross-region calls; that is why we use CDNs, regional deployments and edge termination. See multi-region architecture and CDN and edge.

Round trips matter more than bandwidth for small requests. A new HTTPS connection to a distant server costs several round trips before any data flows. Connection reuse, HTTP/2 and HTTP/3, and terminating TLS near the user save hundreds of milliseconds. See what happens when you type a URL.

Consensus and synchronous replication cost round trips. A strongly consistent write across three regions waits for at least one cross-region round trip. Within one region it is a millisecond or two. See Spanner and distributed SQL.

Sequential beats random. Reading 1 MB sequentially from SSD takes about the same time as a handful of random reads; batch and stream data rather than fetching it piecemeal.

Using them in estimates

Example: a feed page must render in 200 ms for users in Europe, served from a European region.

  • Network to the region and back: about 20 to 40 ms.
  • Budget left for the backend: about 150 ms.
  • A cache hit costs about 1 ms; a database query about 5 ms; a call to another service about 1 to 2 ms plus its own work.
  • Ten sequential database queries (50 ms) fit, but are risky for the tail; parallelise and cache instead. See tail latency.
  • A synchronous call to a US region (about 100 ms round trip) would consume most of the budget: avoid it.

Throughput from latency: one thread doing 5 ms queries sequentially manages about 200 per second; to handle 10,000 per second you need about 50 concurrent queries in flight. See queueing theory and capacity planning.

Throughput numbers worth knowing

ResourceRough capacity
A single Redis instance100,000+ simple operations per second
A well-tuned relational database primarythousands to tens of thousands of transactions per second
A Kafka brokerhundreds of MB per second
A web server instancethousands of simple requests per second
A 10 Gbps network linkabout 1.25 GB per second
NVMe SSDhundreds of thousands of random reads per second; several GB per second sequential

These vary widely with workload, but they let you sanity-check a design in seconds. See back-of-envelope estimation.

Where this shows up in questions

Checklist

  • Memory, SSD, disk and network latencies by order of magnitude.
  • No cross-region calls in user-facing request paths.
  • Minimise round trips: connection reuse, edge termination, batching.
  • Prefer sequential I/O and in-memory hot data.
  • Convert latency into concurrency with Little's law.
  • Sanity-check component throughput with rough capacity numbers.

Open in your browser to sign in

Google does not allow sign-in inside this app's built-in browser. Open this page in Safari and sign in there. The link opens this same page.

Tap the ⋯ or share button at the top or bottom of the screen, then Open in browser. Or copy the link and paste it into Safari.