IoT device fleets and telemetry
Designing backends for millions of connected devices: MQTT and persistent connections, device identity and provisioning, telemetry ingestion and time-series storage, device shadows and desired state, commands to devices, offline buffering, firmware updates, and fleet monitoring and security.
Reading is half of it. See this used in a real interview: walk through Design a Metrics and Monitoring System →
Smart meters, thermostats, scooters, industrial sensors and connected cars all report data and receive commands. An IoT backend looks like a mix of a chat system (millions of persistent connections), a metrics system (endless telemetry) and a device management system (configuration, commands, updates). It comes up in interviews as "design a smart home platform" or "design telemetry for a fleet of scooters", and the patterns are worth knowing.
Connectivity
Devices are often low-power, on flaky cellular or Wi-Fi networks, behind NATs:
- MQTT is the common protocol: lightweight publish/subscribe over a persistent TCP (or WebSocket) connection, with small headers, keepalives, quality-of-service levels (at most once, at least once, exactly once at the protocol level) and retained messages.
- Devices connect outbound to the cloud (they usually cannot accept incoming connections), so commands travel over the device's existing connection.
- MQTT brokers or managed IoT hubs hold millions of connections, clustered and partitioned by device id. See presence and connection management.
- Very constrained devices may use CoAP over UDP or batch uploads over HTTP.
Identity and provisioning
- Each device has a unique identity, typically an X.509 certificate or key stored in secure hardware, used for mutual TLS. Shared passwords across a fleet are a classic disaster.
- Provisioning: at manufacture or first boot, devices enrol, receive credentials and are assigned to an owner or tenant.
- Credentials can be revoked per device (stolen or decommissioned). See security and auth.
Telemetry ingestion
- Devices publish readings to topics like
devices/{id}/telemetry, often batched. - The broker forwards into a stream (Kafka) partitioned by device id. See how Kafka works.
- Consumers write to a time-series store (with downsampling and retention), evaluate alert rules ("temperature above threshold for 5 minutes"), and feed analytics. See time-series data and windowing and watermarks.
- Device clocks drift and data arrives late after offline periods; use event time from the device with sanity checks, and accept out-of-order data.
Estimate: one million devices sending a reading every 10 seconds is 100,000 messages per second, typically small. See estimation worked examples.
Device shadows (digital twins)
Devices go offline, so the cloud keeps a shadow per device:
- Reported state: last values the device reported (firmware version, settings, sensor readings).
- Desired state: what the owner or system wants (target temperature 21 °C, lock engaged).
- When the device connects, it receives the difference and applies it, then reports the new state.
Apps read and write the shadow instead of talking to the device directly, so they work when the device is asleep. Version numbers on shadow updates prevent stale writes. See optimistic vs pessimistic locking.
Commands
- Immediate commands (unlock a scooter) are published to the device's command topic; the device acknowledges execution. Use command ids for idempotency and timeouts for devices that do not respond. See idempotency keys in practice.
- Commands for offline devices either expire or are queued until the next connection, depending on the command.
- Fleet-wide commands are sent in waves, not all at once. See the thundering herd problem.
Offline buffering
Devices buffer telemetry locally when disconnected and upload when back online, oldest first or newest first depending on what matters. Backends must handle bursts of backlog after a regional outage, when thousands of devices reconnect and upload at once: use jittered reconnects and rate limits. See load shedding and backpressure.
Firmware updates
Over-the-air updates must be signed, staged, resumable and safely revertible (A/B partitions), with health reporting after update. See distributing software updates at scale.
Fleet monitoring
Track connected devices, connection churn, message rates, error codes, battery and signal levels, firmware versions and devices that have gone silent (a "last seen" timestamp with alerts). Group by model, region and firmware to spot systematic problems quickly. See observability and operations.
Multi-tenancy and privacy
Consumer IoT data (home presence, location) is sensitive: per-tenant isolation, encryption, minimal retention and clear deletion. Enterprise platforms isolate customers' fleets and data. See multi-tenancy and privacy and data deletion.
In the interview
"Devices authenticate with per-device certificates over MQTT to a clustered broker; telemetry flows into Kafka partitioned by device, then to a time-series store with downsampling and an alerting stream job. Each device has a shadow with reported and desired state; apps update desired state and devices sync on connect. Commands carry ids and expiries; firmware updates are signed and staged; reconnects use jittered backoff." Relevant to scooter or vehicle fleets alongside Design Uber and to Design a Monitoring System.
Checklist
- MQTT (or similar) persistent connections, outbound from devices.
- Per-device certificates, provisioning and revocation.
- Telemetry through a partitioned stream into time-series storage; late data handled.
- Device shadows with reported and desired state, versioned.
- Idempotent, expiring commands; staged fleet-wide actions.
- Offline buffering and jittered reconnects; signed, staged firmware updates.
- Fleet health monitoring; tenant isolation and privacy.