SysDesignPrep.com
Study guide 132 of 183

Trust and safety, moderation and abuse prevention

Designing defences against spam, fake accounts, fraud and harmful content: layered checks at signup, posting and payment, rules plus machine learning, review queues, hashing known-bad media, reputation, rate limits and feedback loops.

Reading is half of it. See this used in a real interview: walk through Design Reddit →

Any system where users create accounts, post content or move money attracts abuse: spam, bots, fake reviews, scams, fraud and harmful content. Interviewers ask "how would you stop spam?" or "how do you handle reported content?" to see whether you design for adversaries, not just honest users. The answer is never one filter; it is layers of cheap and expensive checks, with humans for the hard cases and feedback loops so the system learns.

Where abuse enters

  • Signup: bots creating thousands of fake accounts.
  • Login: credential stuffing and account takeover. See security and auth.
  • Content creation: spam, scams, harassment, illegal or harmful media.
  • Engagement: fake likes, votes, follows and reviews to manipulate ranking.
  • Payments: stolen cards, chargebacks, promotion abuse. See payments and ledgers.
  • Messaging: unsolicited messages and scams in DMs.

Layered defences

Order checks from cheap and fast to expensive and slow:

  1. Rate limits per account, IP, device and network: stop bulk activity outright. See Design a Rate Limiter.
  2. Friction when risk is elevated: CAPTCHAs, phone or email verification, delays for new accounts.
  3. Rules: blocklists, known-bad domains, keyword patterns, velocity rules ("more than 20 posts in a minute from a day-old account").
  4. Hash matching of media against databases of known illegal or banned content (perceptual hashes such as PhotoDNA or PDQ), which survive resizing and small edits.
  5. Machine learning classifiers for text, images and behaviour, producing risk scores.
  6. Human review for uncertain or high-impact cases.

Synchronous versus asynchronous

  • Synchronous checks run before the action succeeds: must be fast (tens of milliseconds) and are reserved for high-confidence signals (rate limits, hash matches, blocklists, a fast model).
  • Asynchronous checks run after posting: heavier models, graph analysis, cross-account correlation. Content may be removed or hidden minutes later.
  • Shadow actions: limiting the reach of suspicious content (visible to the author, not distributed) without telling the abuser, so they do not learn the boundary.

Reputation and graphs

Individual items are hard to judge; actors are easier. Track reputation signals per account, device, IP range and payment instrument: account age, verified contacts, history of reports and removals. Abusers operate in rings, so graph analysis (shared devices, IPs, payment methods, coordinated timing) catches networks of fake accounts that look normal one by one. See graph data.

Reports and review queues

  • Users report content; reports are deduplicated per item and weighted by reporter reliability.
  • A priority queue orders review work by severity, reach (views per hour) and confidence, so a viral harmful post is reviewed before a low-reach borderline one.
  • Reviewers see context, act through tooling with consistent policies, and every decision is logged for appeals and audit.
  • Reviewer wellbeing matters for disturbing content: blurring, limits, rotation.

Feedback loops

Every reviewer decision, appeal outcome and confirmed fraud case becomes training data for the classifiers and evidence for new rules. Measure precision (how often you wrongly block real users) and recall (how much abuse gets through), per category, and watch for adversaries adapting: a sudden drop in detections may mean they found a gap, not that abuse stopped.

Integrity in ranking and counts

Manipulated engagement corrupts feeds and leaderboards. Discount votes and likes from low-reputation or coordinated accounts, detect sudden unnatural spikes, and use reputation-weighted signals in ranking. See Design Reddit and recommendation systems.

In dating and social apps

Dating apps add specific risks: romance scams, catfishing with stolen photos, and harassment. Defences include photo verification (selfie matching), reverse image search for stolen photos, message scanning for scam patterns and off-platform payment requests, and easy blocking and reporting. See Design Tinder and Design Instagram.

In the interview

When users create content or accounts, add a short section: rate limits and friction at signup and posting, fast synchronous checks (hash matching, blocklists, a light model), asynchronous heavier classifiers, reputation and graph signals, a prioritised human review queue, and feedback into the models.

Checklist

  • Rate limits and friction proportional to risk.
  • Synchronous fast checks, asynchronous heavy checks.
  • Perceptual hash matching for known-bad media.
  • Reputation per account, device, IP and payment method; graph analysis for rings.
  • Severity- and reach-prioritised review queues with appeals.
  • Decisions fed back into models; precision and recall tracked.

Open in your browser to sign in

Google does not allow sign-in inside this app's built-in browser. Open this page in Safari and sign in there. The link opens this same page.

Tap the ⋯ or share button at the top or bottom of the screen, then Open in browser. Or copy the link and paste it into Safari.