Compression in system design
Where compression saves money and latency: HTTP compression with gzip, Brotli and Zstandard, compression in Kafka and databases, columnar and time-series encodings, images and video codecs, ratio versus speed trade-offs, dictionary compression, and when compression hurts.
Reading is half of it. See this used in a real interview: walk through Design a Metrics and Monitoring System →
Compression trades CPU for fewer bytes. At scale, fewer bytes means lower bandwidth bills, less storage, faster transfers and more data in memory. Logs often shrink ten times, metrics even more, and the right video codec halves delivery costs. Mentioning compression at the right place in a design (and its costs) is an easy way to show practical experience, especially in storage-heavy and bandwidth-heavy questions.
The basic trade-off
Every algorithm sits somewhere on a curve of ratio (how small) versus speed (how fast to compress and decompress):
| Algorithm | Ratio | Speed | Typical use |
|---|---|---|---|
| LZ4, Snappy | modest | very fast | databases, Kafka, in-memory and RPC payloads |
| gzip (DEFLATE) | good | moderate | HTTP, files, universal compatibility |
| Zstandard (zstd) | good to very good, tunable | fast | Kafka, storage, logs, general purpose default |
| Brotli | very good for text | slow at high levels, fast to decompress | static web assets, HTTP |
| xz, LZMA | very high | slow | archives, cold storage |
Decompression is usually much faster than compression, so content compressed once and read many times (static assets, archived data) can use slower, stronger settings.
On the web
- Compress text responses (HTML, CSS, JavaScript, JSON) with Brotli or gzip based on
Accept-Encoding; typical savings are 70 to 90 % for text. - Precompress static assets at build time at maximum level; compress dynamic responses at a fast level.
- Do not compress already-compressed formats (JPEG, PNG, MP4, zip): it wastes CPU for nothing.
- Use
Vary: Accept-Encodingso caches store the right variants. See HTTP caching.
Messaging and RPC
- Kafka compresses record batches on the producer (zstd, LZ4, Snappy or gzip); brokers store and serve them compressed, saving disk and network across the cluster. Larger batches compress better. See how Kafka works.
- gRPC supports per-message compression; binary formats like Protobuf are already compact, so gains are smaller than with JSON. See schema evolution and serialization.
Databases and storage
- Storage engines compress pages or blocks (LZ4, zstd), trading a little CPU for more data in cache and fewer disk reads. See storage engines.
- Columnar formats compress far better than row formats, because a column's values are similar: run-length encoding, dictionary encoding (store each distinct string once), bit packing and delta encoding before general-purpose compression. See data lakes and lakehouses.
- Time-series databases use specialised encodings: delta-of-delta for timestamps and XOR encoding for floating-point values (from Facebook's Gorilla), often reaching around 1.4 bytes per data point instead of 16. See time-series data and Design a Monitoring System.
Logs
Logs are highly repetitive text and commonly compress 5 to 20 times. Compress on the agent before shipping, store compressed in object storage, and use structured fields so repeated keys compress well. See distributed tracing and structured logging.
Dictionary compression
Small messages (a few hundred bytes) compress poorly on their own, because there is little repetition within one message. Dictionary compression (zstd supports trained dictionaries) shares common patterns across messages, greatly improving ratios for small JSON payloads, cache entries or database rows.
Images and video
Media dominates bandwidth for most consumer products, and its compression is format-specific and lossy:
- Images: WebP and AVIF are much smaller than JPEG at similar quality; resize to the display size first. See media uploads and processing.
- Video: H.264 is universal; HEVC, VP9 and AV1 save roughly 30 to 50 % bandwidth at similar quality, at higher encoding cost. Platforms encode popular videos with expensive codecs and settings because the savings multiply over millions of views. See video streaming and Design YouTube.
Deduplication versus compression
Compression removes redundancy within a stream; deduplication removes identical chunks across files and users (content addressing). Large storage systems use both. See hashing and encoding.
When compression hurts
- CPU-bound services where bandwidth is cheap (inside a data centre with small payloads).
- Latency-critical tiny messages.
- Already-compressed or encrypted data (encryption makes data look random; compress before encrypting, and beware of compression side-channel attacks on secrets mixed with attacker-controlled data in HTTPS).
In the interview
Mention compression where bytes dominate: "logs and events are compressed with zstd in Kafka and stored as compressed Parquet, roughly 10x smaller"; "metrics use delta-of-delta and XOR encoding"; "video is encoded with a modern codec for popular titles to cut CDN egress". For a web crawler, compress stored pages; for a message queue, compress batches at the producer.
Checklist
- Pick algorithms by ratio versus speed; zstd as a strong default.
- Brotli or gzip for text over HTTP; skip already-compressed media.
- Producer-side batch compression in Kafka.
- Columnar and time-series encodings for analytical and metric data.
- Dictionary compression for small messages.
- Modern image and video codecs where bandwidth costs dominate.
- Know when CPU, latency or security make compression a bad idea.