Video streaming: transcoding, HLS, DASH and adaptive bitrate
How online video works end to end: uploads, transcoding into a bitrate ladder, segmenting for HLS and DASH, adaptive bitrate in the player, CDN delivery, live versus on-demand, and the numbers behind it.
Reading is half of it. See this used in a real interview: walk through Design YouTube →
Video is most of the internet’s traffic, and "design YouTube" or "design Netflix" is a staple interview. The core ideas are the same for both: encode each video into several qualities, cut them into small segments, let the player pick a quality segment by segment, and serve everything from a CDN. This guide explains each step at interview depth.
The pipeline at a glance
- Upload the original file (often gigabytes) to object storage, resumably.
- Transcode it into a ladder of resolutions and bitrates.
- Package each rendition into short segments plus a manifest (HLS or DASH).
- Store segments in object storage and serve them through a CDN.
- The player reads the manifest and switches between renditions as bandwidth changes.
Uploads
Large files over flaky connections need multipart, resumable uploads directly to object storage with presigned URLs, so your API servers never carry the bytes and a dropped connection resumes from the last part. See object storage and large files.
Transcoding and the bitrate ladder
The original is re-encoded into a ladder of renditions, for example:
| Rendition | Bitrate (H.264) |
|---|---|
| 240p | 0.3–0.4 Mbit/s |
| 480p | 1–1.5 Mbit/s |
| 720p | 2.5–3 Mbit/s |
| 1080p | 4.5–6 Mbit/s |
| 4K | 15–20 Mbit/s |
Newer codecs (HEVC, VP9, AV1) reach the same quality at 30 to 50 % lower bitrates but cost more CPU to encode, so platforms encode popular videos in more codecs than unpopular ones.
Transcoding is CPU-heavy (often several times the video’s duration on one core per rendition), so it is parallelised: split the video into chunks of a few seconds at keyframes, transcode chunks on many workers at once, and stitch the results. A one-hour video can be ready in minutes. Per-title encoding tunes the ladder to each video’s complexity: a cartoon needs far fewer bits than a grainy action film. See Design YouTube.
Segments and manifests: HLS and DASH
Each rendition is cut into segments of 2 to 6 seconds. A manifest lists the renditions and their segment URLs.
- HLS (HTTP Live Streaming, from Apple):
.m3u8playlists; required on Apple devices. - MPEG-DASH: an XML manifest (
.mpd); widely used elsewhere. - CMAF lets both share the same fMP4 segment files, so you store one set of segments.
Because segments are plain files fetched over HTTP, the whole system rides on ordinary web infrastructure: CDNs, caches, range requests.
Adaptive bitrate (ABR) in the player
The player downloads the manifest, starts with a low or middle rendition so the first frame appears quickly, then picks each next segment’s rendition based on:
- Measured throughput of recent downloads, and
- Buffer level: a full buffer allows a higher quality; a draining buffer forces a lower one.
Buffer-based logic avoids stalls better than throughput estimates alone, which are noisy on mobile. Quality may drop for a few seconds; playback should not stop. Startup time and rebuffering ratio are the metrics viewers feel. See Design Netflix.
Delivery through CDNs
Video is enormous: 40 M concurrent streams at 5 Mbit/s is 200 Tbit/s. Only CDNs can serve it, and for the largest services, caches placed inside ISPs. Segments are immutable (a new encode gets new URLs), so they cache perfectly. Popular content is pushed to edges ahead of demand; the long tail is fetched on a miss. See CDNs and edge computing.
Storage tiering matters too: most views go to a small share of videos, so old, unpopular renditions move to cheaper storage, and some platforms delete rarely watched high-bitrate renditions and re-create them on demand.
Live streaming
Live adds a clock:
- The broadcaster sends a stream (RTMP or SRT) to an ingest server, which transcodes in real time into the ladder.
- Segments are published as they are produced; the manifest is updated continuously.
- Latency: classic HLS with 6-second segments is 15 to 30 seconds behind real time. Low-latency HLS/DASH with partial segments brings it to 2 to 5 seconds; sub-second latency needs WebRTC, at higher cost and lower scale.
- The CDN must handle a sudden audience for one stream: request coalescing at the edge so thousands of viewers asking for the newest segment make one origin request.
DRM and access control
Paid content is encrypted; the player obtains a license (Widevine, FairPlay, PlayReady) from a license server after an entitlement check. Segment URLs are often signed and short-lived so they cannot be hotlinked.
Numbers to remember
- 1080p at 5 Mbit/s ≈ 2.25 GB per hour; 4K at 16 Mbit/s ≈ 7 GB per hour.
- A full ladder across codecs is typically 2 to 4 times the size of the highest single rendition.
- 500 hours of video uploaded per minute (YouTube scale) is about 30,000 hours an hour to transcode.
Checklist
- Resumable direct uploads to object storage.
- Parallel, chunked transcoding into a ladder; codecs chosen by popularity.
- Segments plus HLS/DASH manifests (CMAF to share segments).
- Buffer-based ABR; startup time and rebuffering as the metrics.
- CDN delivery with immutable segment URLs; storage tiering for the tail.
- For live: real-time transcoding, low-latency segments, edge coalescing.
- DRM and signed URLs for paid content.