SysDesignPrep.com
Study guide 37 of 183

Video calls and WebRTC

How Zoom-style video calling works: WebRTC media and signaling, NAT traversal with STUN and TURN, peer-to-peer versus SFU versus MCU, simulcast and adaptive quality, scaling large meetings with cascaded SFUs, recording, latency budgets and reliability.

Reading is half of it. See this used in a real interview: walk through Design WhatsApp →

Video calls have the strictest latency requirements of any common product: conversation feels broken above about 300 to 400 ms of mouth-to-ear delay, so there is no time for buffering, segmenting or retries the way video-on-demand works. "Design Zoom" or adding calls to a chat app is a popular interview extension. The answer centres on WebRTC, media servers and network traversal.

Media versus signaling

  • Media: the audio and video streams, sent over UDP (SRTP) for low latency; losing a packet is better than waiting for it.
  • Signaling: setting up the call (who is calling, which codecs, network addresses), exchanged through your own servers, typically over WebSockets. WebRTC does not define signaling; you design it. See WebSockets vs SSE vs long polling.

A call starts with signaling (offer and answer describing media capabilities), then media flows directly or via media servers.

NAT traversal: STUN and TURN

Most devices sit behind NATs and firewalls, so they cannot simply connect to each other:

  • STUN servers tell a device its public address, enabling direct peer-to-peer connections when NATs allow.
  • TURN servers relay media when direct paths fail (restrictive corporate networks, symmetric NATs). Relaying costs bandwidth, so TURN capacity is a real cost line.
  • ICE tries candidate paths (local, STUN-derived, TURN) and picks the best working one.

Topologies

TopologyHowGood forLimits
Peer-to-peer mesheach participant sends to every other1:1 callsupload grows with participants; breaks beyond a few people
SFU (selective forwarding unit)each participant sends once to a server, which forwards streams to others without decodinggroup calls, the modern defaultserver bandwidth; clients receive many streams
MCU (multipoint control unit)server decodes, mixes into one stream, re-encodesvery weak clients, legacy systemsheavy CPU on servers, added latency

Most products use peer-to-peer (or a TURN relay) for 1:1 calls and SFUs for groups.

Adapting quality

Networks vary widely, so quality must adapt per receiver:

  • Simulcast: each sender encodes several resolutions (for example 180p, 360p, 720p); the SFU forwards the right layer to each receiver based on their bandwidth and layout (thumbnails get low resolution, the active speaker high).
  • SVC (scalable video coding): one stream with layers the SFU can drop.
  • Congestion control and bandwidth estimation adjust bitrates continuously.
  • Audio first: under bad conditions, keep audio and drop video; use forward error correction and packet loss concealment for audio.

Large meetings and webinars

  • An SFU on one machine handles tens to a few hundred participants. Bigger meetings cascade SFUs across servers and regions: each region's SFU receives local senders and exchanges streams with other SFUs.
  • Only a few participants send video at once (active speakers); others are mostly receivers.
  • For thousands of viewers, convert the meeting into a live stream (HLS) with a few seconds of delay. See live streaming.

Placement and routing

Latency is dominated by distance. Place media servers in many regions; connect each participant to the nearest SFU and connect SFUs over good backbone links. Choose the meeting's anchor region by where participants are. See global traffic management and latency numbers.

Features around the call

  • Recording: a bot participant (or the SFU) receives the streams and writes them to object storage, often mixing later. See object storage and files.
  • Chat, reactions and presence over the signaling channel. See presence and connection management.
  • Screen sharing as another video track, optimised for sharp text at lower frame rates.
  • Transcription and captions from audio streams in real time.
  • Encryption: media is always encrypted in transit (DTLS-SRTP); end-to-end encryption through SFUs needs insertable streams so the SFU forwards without decrypting. See end-to-end encryption.

Reliability

  • Reconnect quickly after network changes (ICE restarts) without dropping the call.
  • Fail over to another SFU if one dies, with clients rejoining automatically.
  • Monitor per-call quality: packet loss, jitter, round-trip time, freezes, audio issues, by region and network.

In the interview

"Clients signal over WebSockets to a call service; 1:1 calls go peer-to-peer via ICE with STUN, falling back to TURN relays; group calls use SFUs placed in many regions, with simulcast so each receiver gets the right resolution; large meetings cascade SFUs; webinars beyond a few hundred viewers switch to HLS. Recording is a bot participant writing to object storage."

Checklist

  • Signaling designed separately from media; media over UDP.
  • STUN, TURN and ICE for connectivity; TURN capacity planned.
  • Peer-to-peer for 1:1, SFUs for groups.
  • Simulcast or SVC and congestion control; audio prioritised.
  • Cascaded, regional SFUs; streaming for huge audiences.
  • Recording, captions, screen share and encryption choices.
  • Fast reconnects, SFU failover and call quality metrics.

Open in your browser to sign in

Google does not allow sign-in inside this app's built-in browser. Open this page in Safari and sign in there. The link opens this same page.

Tap the ⋯ or share button at the top or bottom of the screen, then Open in browser. Or copy the link and paste it into Safari.