Video calls and WebRTC
How Zoom-style video calling works: WebRTC media and signaling, NAT traversal with STUN and TURN, peer-to-peer versus SFU versus MCU, simulcast and adaptive quality, scaling large meetings with cascaded SFUs, recording, latency budgets and reliability.
Reading is half of it. See this used in a real interview: walk through Design WhatsApp →
Video calls have the strictest latency requirements of any common product: conversation feels broken above about 300 to 400 ms of mouth-to-ear delay, so there is no time for buffering, segmenting or retries the way video-on-demand works. "Design Zoom" or adding calls to a chat app is a popular interview extension. The answer centres on WebRTC, media servers and network traversal.
Media versus signaling
- Media: the audio and video streams, sent over UDP (SRTP) for low latency; losing a packet is better than waiting for it.
- Signaling: setting up the call (who is calling, which codecs, network addresses), exchanged through your own servers, typically over WebSockets. WebRTC does not define signaling; you design it. See WebSockets vs SSE vs long polling.
A call starts with signaling (offer and answer describing media capabilities), then media flows directly or via media servers.
NAT traversal: STUN and TURN
Most devices sit behind NATs and firewalls, so they cannot simply connect to each other:
- STUN servers tell a device its public address, enabling direct peer-to-peer connections when NATs allow.
- TURN servers relay media when direct paths fail (restrictive corporate networks, symmetric NATs). Relaying costs bandwidth, so TURN capacity is a real cost line.
- ICE tries candidate paths (local, STUN-derived, TURN) and picks the best working one.
Topologies
| Topology | How | Good for | Limits |
|---|---|---|---|
| Peer-to-peer mesh | each participant sends to every other | 1:1 calls | upload grows with participants; breaks beyond a few people |
| SFU (selective forwarding unit) | each participant sends once to a server, which forwards streams to others without decoding | group calls, the modern default | server bandwidth; clients receive many streams |
| MCU (multipoint control unit) | server decodes, mixes into one stream, re-encodes | very weak clients, legacy systems | heavy CPU on servers, added latency |
Most products use peer-to-peer (or a TURN relay) for 1:1 calls and SFUs for groups.
Adapting quality
Networks vary widely, so quality must adapt per receiver:
- Simulcast: each sender encodes several resolutions (for example 180p, 360p, 720p); the SFU forwards the right layer to each receiver based on their bandwidth and layout (thumbnails get low resolution, the active speaker high).
- SVC (scalable video coding): one stream with layers the SFU can drop.
- Congestion control and bandwidth estimation adjust bitrates continuously.
- Audio first: under bad conditions, keep audio and drop video; use forward error correction and packet loss concealment for audio.
Large meetings and webinars
- An SFU on one machine handles tens to a few hundred participants. Bigger meetings cascade SFUs across servers and regions: each region's SFU receives local senders and exchanges streams with other SFUs.
- Only a few participants send video at once (active speakers); others are mostly receivers.
- For thousands of viewers, convert the meeting into a live stream (HLS) with a few seconds of delay. See live streaming.
Placement and routing
Latency is dominated by distance. Place media servers in many regions; connect each participant to the nearest SFU and connect SFUs over good backbone links. Choose the meeting's anchor region by where participants are. See global traffic management and latency numbers.
Features around the call
- Recording: a bot participant (or the SFU) receives the streams and writes them to object storage, often mixing later. See object storage and files.
- Chat, reactions and presence over the signaling channel. See presence and connection management.
- Screen sharing as another video track, optimised for sharp text at lower frame rates.
- Transcription and captions from audio streams in real time.
- Encryption: media is always encrypted in transit (DTLS-SRTP); end-to-end encryption through SFUs needs insertable streams so the SFU forwards without decrypting. See end-to-end encryption.
Reliability
- Reconnect quickly after network changes (ICE restarts) without dropping the call.
- Fail over to another SFU if one dies, with clients rejoining automatically.
- Monitor per-call quality: packet loss, jitter, round-trip time, freezes, audio issues, by region and network.
In the interview
"Clients signal over WebSockets to a call service; 1:1 calls go peer-to-peer via ICE with STUN, falling back to TURN relays; group calls use SFUs placed in many regions, with simulcast so each receiver gets the right resolution; large meetings cascade SFUs; webinars beyond a few hundred viewers switch to HLS. Recording is a bot participant writing to object storage."
Checklist
- Signaling designed separately from media; media over UDP.
- STUN, TURN and ICE for connectivity; TURN capacity planned.
- Peer-to-peer for 1:1, SFUs for groups.
- Simulcast or SVC and congestion control; audio prioritised.
- Cascaded, regional SFUs; streaming for huge audiences.
- Recording, captions, screen share and encryption choices.
- Fast reconnects, SFU failover and call quality metrics.