Real-time fraud detection
Designing fraud and risk systems for payments and marketplaces: the decision point and latency budget, rules plus machine learning, real-time velocity features, device and network signals, graph features, step-up actions, review queues, feedback and label delay, and measuring precision and losses.
Reading is half of it. See this used in a real interview: walk through Design a Payment System →
Every system that moves money attracts fraud: stolen cards, account takeovers, fake listings, promotion abuse, collusion between riders and drivers. A risk system must decide in tens of milliseconds whether to allow, block or challenge an action, and it must keep learning as fraudsters adapt. Payment and marketplace interviews often include "how would you prevent fraud?"; a structured answer stands out.
The decision point
Risk checks sit at key moments: signup, login, adding a payment method, checkout, payout, changing account details. At each, the calling service sends an event to a risk service and waits for a decision:
- Allow.
- Block (with a reason code).
- Step up: require 3-D Secure, a one-time code, ID verification or manual review.
- Allow but hold: let the action proceed, delay payout or fulfilment until review.
The budget is usually 50 to 200 ms, so the decision path must use precomputed features and fast models. If the risk service is down, fail to a safe default per action (often allow small amounts, step up large ones). See rate limiting and resilience.
Rules and models
- Rules are fast to write and explain: block known bad cards, limit five payment attempts per hour, require verification above a threshold for new accounts. Good for known patterns and immediate response to an attack.
- Machine learning models (often gradient-boosted trees) score the probability of fraud from hundreds of features, catching subtler patterns.
Most systems combine them: a model score feeds rules that map scores and context to actions. Rules can be changed in minutes during an attack; models are retrained regularly. See feature flags and A/B testing for safely rolling out rule changes.
Features that matter
| Type | Examples |
|---|---|
| Velocity | attempts per card, account, device or IP in the last minute, hour and day |
| Account history | account age, past chargebacks, typical spend, usual locations |
| Device and network | device fingerprint, emulator detection, IP reputation, proxy or VPN use, distance from billing address |
| Transaction | amount, merchant category, time of day, mismatch between shipping and billing |
| Behaviour | typing speed, navigation pattern, copy-pasted card numbers |
| Graph | shared devices, cards, addresses or phone numbers with known fraudsters |
Velocity features need real-time counters maintained by stream processing and served from a low-latency store. See counting at scale and feature stores and ML serving.
Graph signals
Fraud rings reuse resources. A graph linking accounts, devices, cards, addresses and IPs reveals clusters: a new account sharing a device with ten banned ones is high risk even if its own history is clean. Compute graph features (connected component size, links to known fraud) offline and incrementally, and look up the latest values at decision time. See graph data.
Labels and feedback
Models need labels, and fraud labels arrive late: a chargeback can come weeks after a purchase. Handle this by:
- Training on data old enough for labels to mature, plus faster proxy labels (manual review outcomes, confirmed account takeovers).
- Logging every decision with its features and score for later training and analysis (point-in-time correct).
- Remembering that blocked transactions never get labels; keep a small, controlled sample of allowed borderline cases or use reviewer judgements to avoid blind spots. See bandits and exploration.
Manual review
Uncertain cases go to analysts through a prioritised queue (by amount and score) with tooling showing the account's history, linked accounts and signals. Decisions feed back as labels and into rules. See trust and safety.
Marketplace fraud
Beyond payments:
- Ride-sharing: GPS spoofing, fake trips between colluding riders and drivers, promotion abuse. See Design Uber.
- Rentals: fake listings, off-platform payment scams, account takeovers of hosts. See Design Airbnb.
- Delivery: refund abuse, fake accounts for new-user promotions. See Design DoorDash.
Measuring
- Fraud loss rate (chargebacks and losses as a share of volume).
- False positive rate: good customers blocked or challenged, which costs revenue and trust.
- Review rate and reviewer workload.
- Approval rate overall and by segment.
The goal is a business trade-off, not zero fraud: blocking everything stops fraud and the business.
In the interview
"Checkout calls a risk service with a 100 ms budget; it fetches real-time velocity features from Redis (fed by a stream processor), account and graph features from the online feature store, scores a model, and applies rules mapping score and amount to allow, step up (3-D Secure) or review. Every decision is logged for training; chargebacks arrive later as labels. Rules can be changed instantly during an attack." See Design a Payment System.
Checklist
- Risk decisions at signup, login, payment method changes, checkout and payout.
- Allow, block, step-up and hold actions; safe defaults when the service is down.
- Rules plus models; rules changeable quickly.
- Real-time velocity, device, network, history and graph features.
- Decision logging, delayed labels and controlled exploration.
- Prioritised manual review feeding back into models.
- Balance fraud losses against false positives.