Ad serving and real-time auctions
How an ad system picks and serves an ad in under 100 ms: targeting and candidate retrieval, click prediction, auctions and pricing, budgets and pacing, frequency caps, impression and click tracking, fraud filtering and billing-grade counting.
Reading is half of it. See this used in a real interview: walk through Design an Ad Click Aggregator →
Ads pay for most of the free internet, and the systems behind them combine nearly every hard topic in system design: strict latency budgets, ranking with machine learning, auctions, real-time budget control and counting that must be accurate enough to bill on. "Design an ad click aggregator" is a common interview question, and feed and video questions often include an ad slot. This guide covers the whole path from request to invoice.
The request path
When a page or feed needs an ad, the ad server has roughly 50 to 100 ms to:
- Receive the request with context: user (or device) id, page or app, placement, location, device.
- Retrieve candidates: ads whose targeting matches (audience, location, keywords, interests) and whose campaigns are active with budget left. Inverted indexes over targeting criteria make this fast.
- Predict for each candidate the probability of a click or conversion with a model.
- Run the auction: rank by expected value and choose winners and prices.
- Apply policies: frequency caps, brand safety, competitive exclusions.
- Return the ad with tracking URLs, and log the decision.
This is the same retrieve-then-rank funnel as recommendations. See recommendation systems.
Ranking and auctions
Advertisers bid per click, per impression or per conversion. To compare them, the system computes expected value per impression: for a cost-per-click bid, bid × predicted click-through rate (often adjusted by a quality score). The highest expected values win.
Pricing:
- Second-price (or generalised second price): the winner pays just enough to beat the next ad, which encourages truthful bidding.
- First-price: the winner pays its bid; common in programmatic exchanges, with bidders shading bids down.
Prediction accuracy directly affects revenue and advertiser outcomes, so models are retrained constantly on click and conversion logs, and calibration (predicted probabilities matching observed rates) matters as much as ranking order.
Budgets and pacing
An advertiser with a daily budget of 1,000 does not want it spent by 9 a.m.:
- Pacing spreads spend over the day by throttling participation in auctions (probabilistically) or by lowering bids, based on spend so far versus a target curve.
- Budget tracking needs fast, approximately real-time spend counters across all ad servers. Each server spends from a local allocation that a central service tops up, so servers do not hit a central counter on every request.
- Small overspend is tolerated and usually not charged; large overspend means lost money.
Frequency caps
"Show this ad at most 3 times per user per day" needs per-user, per-campaign counters read on every request: a low-latency key-value store with TTL counters. Approximate is acceptable. See counting at scale.
Tracking impressions and clicks
- An impression is logged when the ad is actually rendered (or viewable), via a pixel or SDK event.
- A click goes through a redirect or a beacon, logged before sending the user to the advertiser.
- Conversions arrive later from advertiser pixels or server-to-server postbacks, and are attributed back to clicks within a window.
Every event carries ids (request, impression, ad, campaign) for joining and deduplication.
Counting for billing
Billing requires numbers that advertisers trust:
- Events flow into a durable log (Kafka), partitioned by ad or campaign. See how Kafka works.
- A stream processor aggregates per ad per minute with deduplication by event id and exactly-once processing, writing results to an OLAP store for dashboards. See delivery semantics.
- A daily batch reconciliation recomputes from the raw log; the batch result is the billing record. See batch and stream processing.
- Late events (mobile devices offline) are handled with watermarks and allowed lateness.
Walk through it in Design an Ad Click Aggregator.
Fraud
Bots and click farms generate fake clicks; publishers may inflate their own impressions. Filter with rules (impossible click rates, data-centre IPs, known bot signatures), anomaly detection and models; mark events invalid before billing, and refund when found later. See trust and safety.
Ads in feeds
In a feed, ads are inserted at positions decided by a blending step that balances revenue with user experience (spacing rules, relevance thresholds, a cap on ad load). See Design a News Feed and Design Instagram.
Privacy
Targeting depends on personal data. Respect consent, limit cross-site tracking, aggregate reporting, and support deletion. See privacy and data deletion.
Checklist
- Under-100 ms path: candidates from targeting indexes, prediction, auction, policies.
- Expected value ranking with calibrated click predictions; second- or first-price rules.
- Pacing and distributed budget allocation.
- Frequency caps with TTL counters.
- Impression, click and conversion events with ids for deduplication and attribution.
- Exactly-once streaming aggregation plus batch reconciliation for billing.
- Invalid traffic filtered before billing.