Recommendation systems
How large recommendation systems are built: candidate generation, ranking and re-ranking, embeddings and two-tower models, features and feature stores, offline and online serving, feedback loops, cold start, and how to evaluate them.
Reading is half of it. See this used in a real interview: walk through Design a News Feed →
Feeds, "videos for you", "people you may know", product suggestions and dating decks are all recommendation systems. In a system design interview you are rarely asked to design the model; you are asked to design the system around it so that it can pick the best 20 items out of millions in under 200 milliseconds and learn from what users do. The architecture is remarkably similar everywhere.
The funnel
Scoring every item with an expensive model for every request is impossible, so recommendations are a funnel:
- Candidate generation (retrieval): from millions of items, cheaply pull a few thousand plausible ones, from several sources at once.
- Ranking: score those thousands with a heavier model that predicts what you care about (click, watch time, like, purchase, match).
- Re-ranking and business rules: diversity, freshness, removing already-seen items, mixing in ads or new creators, policy filters.
Each stage is more expensive per item and handles fewer items. This structure is what lets the whole thing run in a few hundred milliseconds.
Candidate generation
Typical sources, merged:
- Social or follow graph: posts from accounts you follow. See graph data.
- Collaborative filtering: items liked by users similar to you.
- Embedding nearest neighbours: represent users and items as vectors; find items whose vectors are close to yours with an approximate nearest-neighbour (ANN) index. See vector search.
- Content-based: items similar to what you just watched or bought.
- Popular and trending in your region, as a fallback and for new users.
A two-tower model is the common modern recipe: one network embeds the user, another embeds the item, trained so that engaged pairs are close. Item embeddings are precomputed and indexed; the user embedding is computed per request and used to query the ANN index.
Ranking
The ranking model sees rich features for each (user, item) pair: the user’s history and preferences, the item’s age and popularity, context (time, device, location) and cross features (has this user engaged with this creator before). It predicts one or more objectives, combined into a score: for example 0.6 × P(watch > 30 s) + 0.3 × P(like) − 0.5 × P(hide).
Choosing the objective is a product decision with real consequences: optimising clicks rewards clickbait; optimising watch time can reward outrage. Mature systems blend several signals and add explicit penalties.
Features and the feature store
Features come from two places: batch (computed daily from history: a user’s top categories over 90 days) and real-time (computed from streams: what the user clicked in the last five minutes). A feature store serves both with low latency at request time and, crucially, provides the same feature values for training, so the model is trained on what it will see in production. Training/serving skew is one of the most common silent failures. See batch and stream processing.
Serving architecture
- Precompute where possible: for users with stable interests, candidates or even full lists can be computed offline and cached, then refreshed with real-time signals at request time.
- Online path: fetch candidates from each source in parallel, fetch features, rank, re-rank, return. Budget latency per stage (for example 30 ms retrieval, 50 ms features, 60 ms ranking).
- Cache results per user for a short time; a feed refresh need not recompute everything.
- Fallbacks: if ranking times out, return candidates ordered by a simple score; if personalisation is down, return popular items. A degraded page beats an empty one.
Feedback loops and exploration
Recommendations shape the data they learn from: items never shown never get clicks, so the model never learns they are good. Counter this with:
- Exploration: show a small share of uncertain items (new creators, new products) to learn about them.
- Logging what was shown, not just what was clicked, including position, so training can correct for position bias.
- Diversity constraints, so one topic does not take over a user’s feed.
Cold start
- New users: popular and trending items, onboarding questions, context (location, device, referrer), then quick adaptation from the first few interactions.
- New items: content features (text, image embeddings) and a guaranteed small exposure budget to gather signals. See Design Tinder.
Evaluation
- Offline: replay held-out logs and measure ranking quality (precision@k, NDCG, AUC on the prediction targets). Fast, but blind to feedback loops.
- Online: A/B tests on the metrics that matter (sessions, retention, conversions, long-term satisfaction surveys), not only clicks. In two-sided marketplaces, run experiments by region or time window, because the two sides affect each other.
In the interview
Draw the funnel: several candidate sources, a feature store, a ranking service, re-ranking rules, and logging that feeds training. Give the latency budget per stage, the objective being optimised, and how cold start and fallbacks work. Leave model internals at "a two-tower model for retrieval and a gradient-boosted or neural ranker" unless asked.
Checklist
- Candidate generation from several sources, including ANN on embeddings.
- A ranking model with an explicit, defensible objective.
- Re-ranking for diversity, freshness and rules.
- Batch and real-time features from one feature store.
- Latency budget per stage, caching and fallbacks.
- Exploration, logged impressions and position bias.
- Cold start for users and items; offline and online evaluation.