SysDesignPrep.com
Study guide 36 of 183

Live video streaming

How live streams reach millions of viewers: ingest with RTMP, SRT or WebRTC, real-time transcoding, segmenting for HLS and DASH, low-latency HLS, CDN fan-out, WebRTC for interactive streams, latency trade-offs, live chat at scale, recording and reliability.

Reading is half of it. See this used in a real interview: walk through Design YouTube →

Live streaming (Twitch, YouTube Live, sports broadcasts, live shopping) looks like video on demand, but nothing can be prepared in advance: video must be encoded, packaged and delivered within seconds of being captured, and a popular stream can attract millions of viewers in minutes. Interviewers extend video questions with "what if it is live?" to test latency trade-offs, real-time pipelines and fan-out.

The pipeline

  1. Capture and ingest: the broadcaster's software sends a single high-quality stream to the nearest ingest server, usually over RTMP, SRT (more resilient to packet loss) or WebRTC.
  2. Transcode: in real time, encode the stream into a ladder of resolutions and bitrates (1080p, 720p, 480p...) on dedicated machines or GPUs.
  3. Package: cut each rendition into short segments (2 to 6 seconds) and update playlists (HLS) or manifests (DASH).
  4. Distribute: push or pull segments through a CDN to edge servers near viewers.
  5. Play: players fetch the playlist, download segments, and switch bitrate with network conditions (adaptive bitrate). See video streaming.

Latency

Live latency (glass to glass) comes from encoding, segment duration, CDN propagation and the player's buffer:

ApproachTypical latencyScaleUse
Standard HLS or DASH15 to 30 secondshuge, cheap via CDNbroadcasts where delay does not matter
Low-latency HLS or DASH (partial segments, chunked transfer)2 to 5 secondshugesports, streaming with chat
WebRTCunder 1 secondharder; needs media serversvideo calls, auctions, interactive streams

Shorter segments and smaller buffers reduce latency but increase requests and the risk of stalls. Choose based on how interactive the product is.

Fan-out to millions

HTTP-based streaming scales through the CDN: millions of viewers request the same segments, served from edge caches; the origin sees requests only from edge or mid-tier caches. Important details:

  • Request collapsing at the edge and mid-tier caches: when a new segment appears, thousands of simultaneous requests for it become one request to the origin.
  • Playlists change every few seconds, so they have very short cache lifetimes; segments are immutable and cache well. See HTTP caching.
  • Multi-CDN for very large events, with client-side or DNS steering. See global traffic management.

WebRTC at scale

WebRTC sends media peer to peer in theory, but for more than a few participants it goes through SFUs (selective forwarding units): servers that receive each sender's stream and forward it to receivers without re-encoding. Large audiences need cascades of SFUs across regions, which is more expensive than CDN delivery. A common hybrid: WebRTC for the few people on stage, low-latency HLS for the large audience.

Live chat

A popular stream's chat can receive tens of thousands of messages per second, each to be shown to every viewer:

  • Viewers connect to chat gateways over WebSockets; messages are published per stream channel and fanned out by the gateways. See presence and connection management.
  • At extreme scale, sample what each viewer sees (nobody can read 1,000 messages per second), rate-limit senders, and batch deliveries.
  • Moderation in the path: spam filters, slow mode, blocked words. See trust and safety.

The same patterns apply to large channels in chat products. See Design Slack.

Reliability

  • Redundant ingest: broadcasters can send to two ingest points; the pipeline fails over between them.
  • Transcoder failover: health-checked workers, with a standby picking up the stream in seconds.
  • Players tolerate brief gaps by buffering; too small a buffer (for low latency) makes stalls more likely.
  • Monitor per-stream health: ingest bitrate, dropped frames, segment publish delay, viewer rebuffering rate.

Recording and replays

Segments are also written to object storage, so the stream becomes a video-on-demand asset right after it ends (often with re-encoding for better quality). Clips and highlights are cut from the stored segments. See object storage and files.

Cost

Transcoding every stream into a full ladder is expensive. Platforms transcode popular streams fully and give small streams fewer renditions (or pass the source through), adding renditions when viewers arrive. CDN egress dominates for big audiences. See cost-aware system design.

In the interview

For a live extension of Design YouTube or Design Netflix: RTMP or SRT ingest at the edge, real-time transcoding into an ABR ladder, low-latency HLS with short partial segments, CDN fan-out with request collapsing, WebRTC only for interactive participants, chat over WebSockets with sampling at scale, and recording to object storage for replays.

Checklist

  • Ingest close to broadcasters with redundancy.
  • Real-time transcoding ladder sized by popularity.
  • Segment duration and buffers chosen for the latency target.
  • CDN fan-out with request collapsing; short TTL playlists, immutable segments.
  • WebRTC and SFUs only where sub-second latency is required.
  • Chat fan-out with rate limits, sampling and moderation.
  • Recording for replays; per-stream health metrics.

Open in your browser to sign in

Google does not allow sign-in inside this app's built-in browser. Open this page in Safari and sign in there. The link opens this same page.

Tap the ⋯ or share button at the top or bottom of the screen, then Open in browser. Or copy the link and paste it into Safari.