SysDesignPrep.com
Study guide 107 of 183

Feature stores and serving machine learning models

The infrastructure behind production machine learning: offline and online feature stores, training-serving skew, point-in-time correct training data, real-time features from streams, model serving with latency budgets, batching, shadow and canary deployments, and monitoring drift.

Reading is half of it. See this used in a real interview: walk through Design a News Feed →

Feeds, search ranking, ETAs, fraud checks and recommendations all call a machine learning model in the request path, usually with tight latency budgets. The model itself is a small part of the system; the rest is getting the right features to it quickly and correctly, serving it reliably, and noticing when it quietly gets worse. Interviewers for ML-heavy products expect you to describe this infrastructure, not just "call the model".

Features

A feature is an input to a model: a user's number of orders in the last 30 days, a restaurant's average prep time, the distance between rider and driver, the text embedding of a post. Features come from three places:

  • Batch features: computed daily or hourly from the data warehouse (long-term aggregates, user profiles).
  • Streaming features: computed continuously from events (orders in the last 10 minutes, current courier count in an area).
  • Request features: available only at request time (current location, time of day, query text).

The feature store

A feature store manages features consistently for training and serving:

PartStorageUsed for
Offline storedata warehouse or lake, full historybuilding training datasets
Online storelow-latency key-value store (Redis, DynamoDB, Cassandra)fetching current features at inference
Registryfeature definitions, owners, versionsdiscovery and reuse

Feature pipelines (batch and streaming) write to both stores from the same definitions. At request time the model service fetches features for the relevant entities (user, item) from the online store in a few milliseconds. See batch and stream processing.

Training-serving skew

The most common production ML bug: the model sees features computed differently in production than in training (different code, different time windows, missing values handled differently), so it performs worse than offline evaluation suggested. Prevent it by computing features once, from shared definitions, for both paths, and by logging the exact features used at serving time and training on those logs where possible.

Point-in-time correctness

Training examples must use feature values as they were at the time of the event, not as they are now. Joining today's "orders in last 30 days" to a click from three months ago leaks future information into training, inflating offline metrics. Feature stores provide point-in-time joins over the offline history to build correct datasets.

Serving models

  • Latency budget: a ranking call might have 50 ms total, including feature fetches. Fetch features in parallel and batch-score many candidates in one call.
  • Batching requests on GPUs or CPUs raises throughput; dynamic batching waits a few milliseconds to fill a batch.
  • Model size: distill or quantise large models; use a cheap model for candidate generation and a heavier one for the final ranking of a few hundred items. See recommendation systems.
  • Fallbacks: if the model or feature store is slow, serve a simpler model, cached scores or a heuristic rather than failing. See rate limiting and resilience.
  • Precompute where possible: nightly scores for items that do not depend on request context.

Deploying models safely

  • Model registry with versions, training data references and evaluation results.
  • Shadow deployment: the new model scores live traffic without affecting results; compare outputs and latency.
  • Canary and A/B tests: route a percentage of traffic, measure online business metrics, not just offline accuracy. See feature flags and A/B testing.
  • Fast rollback to the previous version.

Monitoring

Models degrade silently as the world changes:

  • Data drift: input feature distributions shift (a new user segment, a broken upstream pipeline producing zeros).
  • Prediction drift: output distributions shift.
  • Performance: actual outcomes (clicks, delivery times) compared with predictions, often with a delay until labels arrive.
  • Feature freshness: streaming features that stop updating.

Alert on these like any other SLI, and retrain on a schedule or when drift is detected. See observability and operations.

Examples

  • Feed ranking: user and post features from the online store, a ranking model scoring hundreds of candidates within the budget. See Design a News Feed.
  • Delivery ETA: streaming features (current kitchen load, courier supply) plus batch features (historical prep time). See Design DoorDash.
  • Dating recommendations: two-sided features and predicted mutual interest. See Design Tinder.

Checklist

  • Batch, streaming and request-time features from shared definitions.
  • Offline and online stores fed by the same pipelines.
  • Logged serving features; point-in-time correct training data.
  • Parallel feature fetches, batch scoring, multi-stage models, fallbacks.
  • Registry, shadow, canary and A/B deployment with rollback.
  • Drift, freshness and outcome monitoring with retraining.

Open in your browser to sign in

Google does not allow sign-in inside this app's built-in browser. Open this page in Safari and sign in there. The link opens this same page.

Tap the ⋯ or share button at the top or bottom of the screen, then Open in browser. Or copy the link and paste it into Safari.