SysDesignPrep.com
Study guide 61 of 183

Location tracking at scale

Ingesting and serving live GPS locations for drivers, couriers and devices: update frequency and battery, the ingest path, latest-location stores versus history, geospatial indexing of moving objects, sharing a live location with riders, map matching, geofencing, privacy and retention.

Reading is half of it. See this used in a real interview: walk through Design Uber →

Ride-hailing, delivery, logistics and social apps track millions of moving devices in real time. Every few seconds each driver's phone reports its position; the system must update a live index for matching, stream positions to the rider watching the car approach, store history for trips and analytics, and do it all without draining batteries. This is a core part of Uber, DoorDash and maps interviews.

How often to update

Update frequency trades accuracy against battery, data and server load:

  • Every 4 to 5 seconds while a driver is available or on a trip is typical.
  • Less often (or only on significant movement) when idle or stationary.
  • More often near pickup or for turn-by-turn navigation.

Phones batch several fixes per upload when connectivity is poor. The client sends position, accuracy, speed, heading and a device timestamp. See mobile system design.

Estimating load

One million active drivers at one update every 4 seconds is 250,000 writes per second, each about 100 bytes, so about 25 MB/s, roughly 2 TB per day of raw history. Writes are small and constant; reads come from matching (nearby queries) and riders watching trips. See estimation worked examples.

The ingest path

  1. Devices send updates over a persistent connection (or HTTP batches) to location gateways. See presence and connection management.
  2. Gateways validate and authenticate, then publish to a stream partitioned by driver id (or by city). See how Kafka works.
  3. Consumers:

- update the latest-location store and the geospatial index; - forward positions to subscribers (the rider on that trip); - write history to a time-series or wide-column store; - feed traffic estimation and analytics.

Separating these consumers keeps the hot path (latest location) fast regardless of history or analytics load.

Latest location and the moving index

The current position is all matching needs:

  • Keep it in memory, keyed by driver id (Redis or an in-process store per city shard), overwritten on each update; no need for durability beyond a few seconds.
  • Maintain a geospatial index by cell (geohash, S2 or H3): when a driver moves to a new cell, remove them from the old cell's set and add them to the new one. Most updates stay in the same cell, so index changes are rarer than location writes. See geohash vs quadtree vs H3.
  • Expire drivers who stop reporting (TTL) so ghosts do not get matched.
  • Partition by city or region; dense cities split further. See dispatch and marketplace matching.

Live location sharing

The rider watching the driver approach subscribes to that trip's channel; the location consumer pushes each new position over the rider's WebSocket. Clients interpolate between updates (animating along the road) so movement looks smooth despite 4-second gaps. Share only during the trip, then stop. See WebSockets vs SSE vs long polling.

History and trips

Geofencing

Many features depend on entering or leaving areas: airport queues, delivery zones, "driver arrived" notifications. Check each update against geofences indexed by cell (cover each fence with cells, then do exact point-in-polygon tests on candidates). Debounce boundary jitter so a driver sitting on a boundary does not flap in and out.

Accuracy and spoofing

GPS is noisy in cities and can be faked. Filter impossible jumps (speed checks), weight by reported accuracy, combine with network and sensor data, and flag suspected spoofing for fraud review. See fraud detection.

Privacy

Location history is sensitive personal data:

  • Collect only while needed (on shift, during trips, while sharing).
  • Short retention for raw points; coarsen or aggregate older data.
  • Restrict access, audit it, and honour deletion requests. See privacy and data deletion.

In the interview

"Drivers send positions every 4 seconds over a persistent connection to gateways, which publish to Kafka partitioned by city. One consumer updates an in-memory latest-location map and H3 cell index per city (with TTLs); another streams positions to riders on active trips; another writes history to a wide-column store partitioned by trip, later map-matched for fares. About 250,000 writes per second for a million drivers." See Design Uber and Design DoorDash.

Checklist

  • Update frequency tuned for battery, accuracy and state (idle, on trip).
  • Load estimated: writes per second, bytes per day.
  • Gateways, a partitioned stream and separate consumers per purpose.
  • In-memory latest locations and a cell index updated on cell changes, with TTLs.
  • Live sharing via WebSockets with client interpolation.
  • History in a write-optimised store; map matching for trips.
  • Geofences with debouncing; spoofing checks; strict privacy and retention.

Open in your browser to sign in

Google does not allow sign-in inside this app's built-in browser. Open this page in Safari and sign in there. The link opens this same page.

Tap the ⋯ or share button at the top or bottom of the screen, then Open in browser. Or copy the link and paste it into Safari.