SysDesignPrep.com
Study guide 162 of 183

Real-time fraud detection

Designing fraud and risk systems for payments and marketplaces: the decision point and latency budget, rules plus machine learning, real-time velocity features, device and network signals, graph features, step-up actions, review queues, feedback and label delay, and measuring precision and losses.

Reading is half of it. See this used in a real interview: walk through Design a Payment System →

Every system that moves money attracts fraud: stolen cards, account takeovers, fake listings, promotion abuse, collusion between riders and drivers. A risk system must decide in tens of milliseconds whether to allow, block or challenge an action, and it must keep learning as fraudsters adapt. Payment and marketplace interviews often include "how would you prevent fraud?"; a structured answer stands out.

The decision point

Risk checks sit at key moments: signup, login, adding a payment method, checkout, payout, changing account details. At each, the calling service sends an event to a risk service and waits for a decision:

  • Allow.
  • Block (with a reason code).
  • Step up: require 3-D Secure, a one-time code, ID verification or manual review.
  • Allow but hold: let the action proceed, delay payout or fulfilment until review.

The budget is usually 50 to 200 ms, so the decision path must use precomputed features and fast models. If the risk service is down, fail to a safe default per action (often allow small amounts, step up large ones). See rate limiting and resilience.

Rules and models

  • Rules are fast to write and explain: block known bad cards, limit five payment attempts per hour, require verification above a threshold for new accounts. Good for known patterns and immediate response to an attack.
  • Machine learning models (often gradient-boosted trees) score the probability of fraud from hundreds of features, catching subtler patterns.

Most systems combine them: a model score feeds rules that map scores and context to actions. Rules can be changed in minutes during an attack; models are retrained regularly. See feature flags and A/B testing for safely rolling out rule changes.

Features that matter

TypeExamples
Velocityattempts per card, account, device or IP in the last minute, hour and day
Account historyaccount age, past chargebacks, typical spend, usual locations
Device and networkdevice fingerprint, emulator detection, IP reputation, proxy or VPN use, distance from billing address
Transactionamount, merchant category, time of day, mismatch between shipping and billing
Behaviourtyping speed, navigation pattern, copy-pasted card numbers
Graphshared devices, cards, addresses or phone numbers with known fraudsters

Velocity features need real-time counters maintained by stream processing and served from a low-latency store. See counting at scale and feature stores and ML serving.

Graph signals

Fraud rings reuse resources. A graph linking accounts, devices, cards, addresses and IPs reveals clusters: a new account sharing a device with ten banned ones is high risk even if its own history is clean. Compute graph features (connected component size, links to known fraud) offline and incrementally, and look up the latest values at decision time. See graph data.

Labels and feedback

Models need labels, and fraud labels arrive late: a chargeback can come weeks after a purchase. Handle this by:

  • Training on data old enough for labels to mature, plus faster proxy labels (manual review outcomes, confirmed account takeovers).
  • Logging every decision with its features and score for later training and analysis (point-in-time correct).
  • Remembering that blocked transactions never get labels; keep a small, controlled sample of allowed borderline cases or use reviewer judgements to avoid blind spots. See bandits and exploration.

Manual review

Uncertain cases go to analysts through a prioritised queue (by amount and score) with tooling showing the account's history, linked accounts and signals. Decisions feed back as labels and into rules. See trust and safety.

Marketplace fraud

Beyond payments:

  • Ride-sharing: GPS spoofing, fake trips between colluding riders and drivers, promotion abuse. See Design Uber.
  • Rentals: fake listings, off-platform payment scams, account takeovers of hosts. See Design Airbnb.
  • Delivery: refund abuse, fake accounts for new-user promotions. See Design DoorDash.

Measuring

  • Fraud loss rate (chargebacks and losses as a share of volume).
  • False positive rate: good customers blocked or challenged, which costs revenue and trust.
  • Review rate and reviewer workload.
  • Approval rate overall and by segment.

The goal is a business trade-off, not zero fraud: blocking everything stops fraud and the business.

In the interview

"Checkout calls a risk service with a 100 ms budget; it fetches real-time velocity features from Redis (fed by a stream processor), account and graph features from the online feature store, scores a model, and applies rules mapping score and amount to allow, step up (3-D Secure) or review. Every decision is logged for training; chargebacks arrive later as labels. Rules can be changed instantly during an attack." See Design a Payment System.

Checklist

  • Risk decisions at signup, login, payment method changes, checkout and payout.
  • Allow, block, step-up and hold actions; safe defaults when the service is down.
  • Rules plus models; rules changeable quickly.
  • Real-time velocity, device, network, history and graph features.
  • Decision logging, delayed labels and controlled exploration.
  • Prioritised manual review feeding back into models.
  • Balance fraud losses against false positives.

Open in your browser to sign in

Google does not allow sign-in inside this app's built-in browser. Open this page in Safari and sign in there. The link opens this same page.

Tap the ⋯ or share button at the top or bottom of the screen, then Open in browser. Or copy the link and paste it into Safari.