Distributing software updates at scale
How apps, games, devices and package managers deliver updates to millions of clients: update manifests and version checks, CDN and peer-assisted delivery, delta patches, signing and integrity, staged rollouts and kill switches, rollback, bandwidth peaks on release day, and offline or constrained devices.
Reading is half of it. See this used in a real interview: walk through Design Netflix →
Shipping a new version to millions of phones, desktops, game consoles, cars or IoT devices is a distribution problem with real risks: a release day can generate more traffic than anything else the company serves, a bad update can brick devices, and a compromised update channel is one of the most damaging attacks possible. Interviews about desktop clients, games or device fleets touch on this, and it combines CDNs, delta encoding, signing and staged rollouts.
The update flow
- The client periodically (with jitter) asks an update service: "I am version 5.2 on platform X, channel stable; is there something newer for me?"
- The service returns a manifest: target version, file list or patch, sizes, hashes, signature, and rollout rules.
- The client downloads from a CDN, verifies hashes and signatures, installs (often in the background), and reports success or failure.
The update service is small and dynamic; the heavy bytes come from the CDN. See CDN and edge.
Integrity and signing
Updates run with high privileges on millions of machines, so the channel must be trustworthy:
- Sign manifests and packages with keys kept in hardware security modules; clients verify against embedded public keys before installing. See encryption and key management.
- Verify hashes of every downloaded file.
- Protect against rollback attacks (serving an old, vulnerable signed version) with version checks and signed metadata expiry; frameworks such as The Update Framework (TUF) formalise this.
- Separate signing keys by role and rotate them; a stolen key should not compromise everything.
Delta updates
Full downloads waste bandwidth when only part changed:
- Binary diffs between consecutive versions (bsdiff-style) shrink patches dramatically for executables.
- Chunk-based updates download only changed chunks, identified by content hashes, which also lets clients skip versions. See delta sync.
- Precompute patches for the most common source versions; fall back to full downloads for very old ones.
Release day traffic
A popular game or OS update can push tens of terabits per second:
- Pre-position content in CDNs (and in ISP-embedded caches) before release.
- Stagger availability by region or randomised cohort to flatten the peak.
- Peer-assisted delivery (devices on the same network sharing pieces, or peer-to-peer swarms) reduces CDN load, used by some game platforms and operating systems.
- Clients back off when servers signal overload. See the thundering herd problem.
Staged rollouts
Never ship to everyone at once:
- Release to internal users, then 1 %, 5 %, 20 %, 100 %, watching crash rates, install failures and key metrics per version. See deployment strategies.
- The update service decides eligibility per client (hash of device id against the rollout percentage, plus targeting by platform, region or hardware).
- A kill switch halts the rollout instantly by changing the manifest.
- Channels (stable, beta, nightly) let enthusiasts test early.
Rollback and recovery
- Keep the previous version installed until the new one starts successfully (A/B partitions on devices and operating systems), and switch back automatically if it fails health checks on boot.
- Server-side, roll back by pointing the manifest at the previous version, though clients that already updated may need a new forward fix.
- Make updates atomic: never leave a device half-updated.
Constrained and offline devices
IoT devices, cars and phones on metered connections:
- Download only on Wi-Fi or while charging when possible; respect user settings. See mobile system design.
- Resumable downloads with range requests and chunk checksums.
- Devices offline for months may need multi-step upgrades through required intermediate versions.
- Fleets report versions and health, so operators can see adoption and stragglers.
App stores
Mobile apps mostly update through app stores, which handle distribution and signing but add review delays and slower rollbacks. That is why apps rely heavily on server-driven configuration and feature flags to change behaviour without a new release. See feature flags and A/B testing.
In the interview
"Clients check a lightweight update service with jitter; it returns a signed manifest chosen by rollout rules. Packages and chunk-level deltas are served from a CDN, pre-positioned before release and staggered by cohort. Clients verify signatures and hashes, install atomically, keep the previous version for automatic rollback, and report health; the rollout advances from 1 % to 100 % on crash-rate gates, with a kill switch." Similar delivery ideas apply to large media catalogues in Design Netflix and sync clients in Design Dropbox.
Checklist
- Update service for manifests; CDN for bytes.
- Signed manifests and packages, hash verification, rollback-attack protection.
- Binary or chunk-level deltas with full-download fallback.
- Pre-positioning, staggering and peer assistance for release peaks.
- Staged rollouts with health gates and a kill switch.
- Atomic installs, automatic rollback, resumable downloads for constrained devices.