Design a Project-to-Contractor Matching System
Run this for someone else. You hold the answers; they do not. Read the prompt, keep the clock, and use the probes below when an answer is thin. Do not show them this page.
Open with this
Match posted projects with qualified, available contractors across a 500 k-person talent pool, and learn from who actually gets hired. Take a couple of minutes on requirements, then we will do some numbers, then the design. I will interrupt to keep us moving.
The clock
- 4 min: functional requirements and scope
- 4 min: non-functional requirements, with numbers
- 5 min: back-of-envelope estimates
- 16 min: high-level design and one or two flows
- 16 min: deep dives and the close
Move them on out loud when a section overruns. The commonest failure is spending twenty minutes on requirements and never reaching a deep dive, and preventing that is your job as much as theirs.
Requirements · 8 min
Listen for: a scoped set of capabilities, an explicit out-of-scope list, and numeric targets rather than adjectives. Prompt with “what are you not building?” if they never scope, and “what number would make that requirement real?” if they say “fast” or “highly available”.
Functional (8)
- Project owners post openings with requirements: Skills, seniority, start date, hours per week, duration, timezone window, budget ceiling, and legal eligibility (work authorization, country, background check). Each requirement is marked must-have or nice-to-have.
- Contractors maintain a profile, availability and preferences: Verified skills, rate, weekly hours free, preferred project types, and the maximum number of concurrent engagements they will take.
- Owners get a ranked shortlist per opening: Top 20 candidates, each with reason codes ("4 of 5 required skills verified, 6 h timezone overlap"). This is a recommendation list, not automatic assignment: a human picks.
- Contractors get a feed of recommended projects: The same match, seen from the other side. A match only happens if both sides say yes, so the contractor feed is half of the product.
- Invite, accept and decline, with reasons: Owner invites, contractor accepts or declines within 48 h. Decline reasons ("rate too low", "not interested in the domain") are both a label and a preference update.
- Respect contractor capacity: A contractor at their concurrent-engagement limit must stop receiving invites, and one with several pending invites must not be shown to everyone.
- Learn from outcomes: Shown, invited, accepted, hired, and project success weeks later. The ranker improves from these without only reinforcing whoever it already showed.
- Out of scope: Payments and invoicing, the vetting interview itself, contracts, messaging, and fraud detection. Vetting results arrive as features.
Non-functional (7)
- Scale (2 M contractors · 500 k active · 15 k openings/day): Request rate is small. The difficulty is in what to show, not how fast.
- Shortlist latency (p99 < 800 ms): An owner waits on a page load. Retrieval, scoring 1 k candidates and re-ranking must fit together.
- Eligibility correctness (0 ineligible candidates shown): Showing someone without work authorization for the country, or already at capacity, is a real harm and a legal problem. This is the requirement that drives the hardest tradeoff: hard rules must never be traded off against a learned score.
- Freshness (availability change visible < 1 min): A contractor who just accepted their second project and hit their limit must drop out of shortlists quickly, or owners invite someone who cannot say yes.
- Match quality (invite → accept > 40 % · hire → success > 80 %): Measured per cohort. Accept rate is the two-sided signal; success is the one that matters but arrives weeks late.
- Fair exposure (new contractors get ≥ 5 % of impressions): A marketplace that only shows proven people starves its supply. Concentration is monitored like latency.
- Explainability and audit (every impression reproducible): Both sides need a reason for a recommendation, and an audit must be able to replay why a person was or was not shown. Log model version, features and score with each impression.
Estimates · 5 min
Ask for two or three numbers, not all of them. What matters is whether they state assumptions, round sensibly, and say what the number implies. Push once with “where did that come from?”
The numbers (7)
- Ranking requests: ~14/s avg · ~55/s peak. 500 k active contractors × 2 feed loads a day = 1 M, plus 200 k owner shortlist loads = 1.2 M/day ÷ 86 400 ≈ 14/s, ×4 at peak. Throughput is not the hard part.
- Candidates after hard filters: ~2 k per opening. 500 k active × ~5 % with the required skills ≈ 25 k; availability, timezone, rate and eligibility keep ~1 in 12 ≈ 2 k. Retrieval hands the top 1 k to the ranker.
- Scoring work: ~55 k scores/s peak. 55 requests/s × 1 k candidates = 55 k model evaluations/s. A gradient-boosted model scores ~100 k rows/s per core, so a few cores. Feature fetches (1 k keys per request) dominate the latency, not the model.
- Embedding index: ~2 GB. 2 M contractors × 256-dim float32 vector = 2 M × 1 KB = 2 GB. One in-memory ANN index per replica, rebuilt nightly and patched on profile change.
- Interaction events: ~25 M/day · ~5 TB/yr. 1.2 M ranking requests × 20 shown = 24 M impressions/day, plus clicks, invites and outcomes ≈ 25 M events × ~500 B (with score, position, propensity, model version) ≈ 12 GB/day ≈ ~5 TB/yr.
- Concurrent engagements: ~450 k. 15 k openings/day × ~30-day average engagement ≈ 450 k live engagements. With a typical limit of 2 each, that needs at least 225 k distinct contractors, 45 % of the active pool. Capacity binds: the same top 5 % cannot fill the market.
- Labels per month: ~450 k hires · success at +30 days. 15 k hires/day × 30 = 450 k hire labels a month against ~720 M impressions: a positive rate of ~0.06 %. Success labels arrive a month later. So the ranker also trains on nearer signals (invite, accept) and weights them.
High-level design · 16 min
Let them draw. Interrupt only to ask what backs a component or what a box actually does. Then pick one flow below and ask them to walk it end to end.
Components (15)
- Owner app: Where a project owner posts openings, reads the ranked shortlist with reason codes, relaxes constraints when the list is thin, and sends invites.
- Contractor app: Profile, availability and preferences, a feed of recommended projects, and the inbox where invites are accepted or declined with a reason. The contractor side is half the match: a list the contractor would decline is a bad list.
- Email / push (SES · APNs · FCM): Delivers invites and reminders. Invites expire after 48 h, so delivery latency and the reminder at 24 h directly move the accept rate.
- Profile service (skill taxonomy): Owns contractor profiles and project specs. Normalises free-text skills onto a taxonomy ("ReactJS", "React.js" → react) and records which skills are verified by assessment. Separates must-have from nice-to-have on every project requirement at write time, so the matcher never has to guess.
- API gateway (auth · rate limits): Authenticates owners and contractors, routes shortlist and feed requests to the match service and invite actions to the allocation service. Stamps a request id that follows the impression into the logs.
- Core DB (Postgres): Source of truth for contractors, projects, openings, invites, engagements and outcomes. Change data capture streams every row change onto the event log, which feeds both the search index and the warehouse.
- Match service (stateless orchestrator): Runs one request end to end: compile the opening into hard filters, retrieve, score, re-rank for capacity and exposure, attach reason codes, and log the impression with its propensity. Hard constraints are applied here as filters and are never inputs the model can outweigh.
- Exposure & capacity (re-ranker · invites): Takes the ranked list and makes it safe for the marketplace: removes anyone at capacity, demotes contractors with many pending invites, caps daily exposure, and reserves exploration slots. Also owns the invite lifecycle and the capacity holds that go with it.
- Event log (Kafka): Impressions (with position, score, propensity and model version), clicks, invites, responses and CDC changes from the core DB. Keyed by contractor id so a contractor's availability changes apply in order.
- Candidate index (OpenSearch · filters + kNN): One document per contractor: taxonomy skills, availability, timezone, rate, eligibility flags, open capacity and a profile embedding. Serves filtered retrieval: hard constraints as filter clauses, then lexical and vector similarity inside the survivors. Updated from CDC within seconds.
- Ranker (GBDT, multi-task): Scores ~1 k candidates with separate heads for P(owner invites), P(contractor accepts) and P(success | hired), combined into one expected-value score. Returns the per-feature contributions that become reason codes.
- Capacity ledger (Redis · slots + holds): Per contractor: max concurrent engagements, active count, pending invite holds with TTL, and impressions today. Every accept is a compare-and-set here so two projects cannot both take the last slot.
- Warehouse (S3 + Iceberg): Every impression and outcome, joined by impression id, with delayed labels filled in as they arrive (accept in days, success in a month). The training set and the audit trail are the same tables.
- Feature store (online Redis · offline Iceberg): Contractor features (verified skills, response rate, past success, rate percentile), project features and pair features (skill overlap, timezone overlap). Online and offline values come from the same definitions, so training and serving see the same numbers.
- Training & eval (daily · counterfactual): Retrains the ranker from logged impressions weighted by inverse propensity, evaluates new models counterfactually on exploration traffic before any A/B test, and runs the concentration and cohort-exposure monitors.
Flows to ask them to walk (5)
- Owner opens a ranked shortlist: Filter, retrieve, score, re-rank, explain, log. The order of those steps is the design: eligibility comes first and is never part of the learned score.
- Owner opens the shortlist for an opening.
- Match service loads the opening and compiles its must-haves into filters.
- Candidate index returns the top 1 k eligible contractors by relevance.
- Ranker scores the 1 k with features and combines its heads into one expected value.
- Exposure & capacity re-ranks the list for the marketplace.
- Match service attaches reason codes and returns the top 20.
- The impression is logged with positions, scores, propensities and model version.
- Invite, accept, and the capacity race: A match needs both sides to agree, and a contractor has a finite number of slots. Holds and a compare-and-set keep the capacity promise honest.
- Owner invites three contractors from the shortlist.
- The invite is stored and the contractor is notified.
- Contractor accepts.
- Allocation commits the engagement only if a slot is still free.
- The capacity change reaches the index, so the contractor stops appearing where they cannot say yes.
- Another contractor declines with a reason; the reason is logged as a label and a preference.
- A contractor who declined for lack of time updates their availability.
- A star contractor floods every shortlist: The best-scored profiles win every ranking, collect dozens of invites they cannot accept, and starve everyone else. The fix lives in the re-ranker, not the model.
- Monitoring shows the top 1 % of contractors receiving 35 % of impressions and invites.
- The re-ranker compares each contractor's pending invites with their open slots.
- A daily exposure budget spreads impressions in proportion to capacity.
- When the star's slots fill, the capacity change removes them from retrieval entirely.
- The quality cost of spreading is measured, not assumed.
- No one meets every hard constraint: The honest answer is an empty list with a map of which constraint to relax, not a list that quietly ignores a must-have. Plus the ranker being down.
- Retrieval with every must-have returns 3 candidates.
- Match service runs count queries with each negotiable must-have relaxed in turn.
- Owner sees the 3 matches plus the tradeoffs that would widen the list.
- The thin result is logged as unmet demand for supply planning.
- Separately, the ranker times out on a request.
- The fallback impression is logged with its own policy name.
- Learning from outcomes without exposure bias: Only shown contractors can be hired, so naive training teaches the model to prefer whoever it already showed. Propensities, exploration and delayed labels fix that.
- Impressions, responses and engagement outcomes land in the warehouse.
- Delayed labels fill in as they arrive.
- Training weights each shown example by the inverse of its propensity.
- A candidate model is evaluated offline on the exploration traffic.
- Cohort and concentration checks gate the release.
- The model ships to the ranker in shadow, then to an opening-level A/B.
Deep dives · 16 min
Pick two. Ask the headline question, let them answer, then use the follow-ups. The follow-ups are where the level gets decided, so leave time for at least three of them.
Keep eligibility out of the learned score
Ask: Why not just give the model timezone, rate and work authorization as features and let it learn what matters?
Good answers name: Hard filters before retrieval; preferences as ranking features, Everything as features in one learned score, Learned score with large penalty features for violations, Post-ranking filter that drops violators from the top 20.
Our pick: Classify every requirement at write time: legal eligibility and certification are always hard, capacity is always hard, and the owner marks the rest as must-have or nice-to-have in the posting form. Must-haves compile to filter clauses in the retrieval query so the ranker never sees an ineligible candidate; nice-to-haves become features. Fast-changing hard state (open slots) is checked again at re-rank time against the capacity ledger because the index can lag by seconds. When the filters leave too few candidates, the system reports which negotiable must-have to relax and how many people that adds, and never relaxes one on its own. LinkedIn's published Recruiter search work follows the same split: structured filters first, learned relevance inside them.
- An owner marks 9 of 10 requirements as must-have and gets zero candidates. What now?
Show zero honestly, then the relaxation map: which single constraint, relaxed by how much, adds how many eligible people. Most owners pick one. Over time, track which must-haves owners relax most often; the posting form can then pre-suggest them as nice-to-have. The system never overrules the owner, but it can make the cost of strictness visible. - A contractor's availability in the index says 30 h/week but they are actually fully booked elsewhere off-platform. How do you catch it?
You cannot see off-platform work, so rely on the contractor: ask for availability confirmation every 2 weeks and decay "available" to unknown if they do not respond. Declines with reason "no capacity" flip it immediately. And treat stale availability as a ranking feature (a confidence), not the hard filter itself: the hard filter is "not known to be unavailable". - Which constraints would you let the model learn to soften, if any?
Only ones the owner did not mark: a rate slightly over the soft target, a secondary skill, seniority. The model can learn that owners frequently hire a mid-level engineer when they asked for senior, and rank accordingly among eligible candidates. It can never move a candidate across a must-have boundary. - How do you prove to an auditor that nobody ineligible was shown?
Each impression log carries the compiled filter and the snapshot version of each shown contractor's eligibility fields. A daily job re-evaluates the filter against the logged values and alerts on any violation. Because the filter is plain code, the auditor can read it; you could not make that argument about a model weight.
Retrieval then ranking
Ask: You have 500 k active contractors. Why not just score all of them with your best model?
Good answers name: Filtered hybrid retrieval (taxonomy + embedding) then a multi-task ranker, Score every active contractor per request, Hand-weighted rule score (skill overlap, rating, response rate), LLM re-ranking of the top 50 against the project description.
Our pick: Retrieval is one OpenSearch query: must-have filter clauses, then BM25 on taxonomy skills and titles plus kNN on a 256-dim profile embedding over the filtered set, merged by reciprocal rank fusion, top 1 k. Recall is measured offline weekly: of the contractors who were hired in the last month, what fraction would the retrieval stage have returned for that opening; target above 95 %. The ranker is a gradient-boosted multi-task model with three heads (P(invite), P(accept | invite), E[success | hired]) whose product is the score; GBDT because the features are tabular, it trains in minutes, and it gives per-feature contributions for reason codes. An embedding two-tower model is a natural next step once there is enough data, used as a retrieval source rather than a replacement for the ranker.
- How would you know retrieval is dropping good candidates?
Retrieval recall at 1 k against a proxy of "good": contractors who were later hired for similar openings, or who were invited from exploration slots and accepted. If recall drops below target for a skill cluster, the usual cause is taxonomy drift (a new framework name) or an embedding that has not seen the new vocabulary. Also watch the share of hires that came from outside the main retrieval path, such as direct search. - Why multiply the heads instead of training one model on "hired"?
Hired is a rare, late label that mixes two decisions: the owner's and the contractor's. Separate heads train on denser labels (invite and accept happen within days), each is easier to diagnose, and the product lets you change the objective (weight success more) without retraining. The cost is that errors compound, so calibration of each head matters. - The ranker is 10x slower after adding pair features. What do you do?
First cut the candidate count: score 1 k with a light first-stage model, then 200 with the heavy one. Precompute the pair features that do not depend on the request (contractor x owner history) nightly. And batch the feature store reads, which are usually the real cost, not the trees. - A contractor asks why they never appear for Python projects. What can you tell them?
Which hard filters they fail most often across recent Python openings (for example "hours available below requested" in 60 % of them), and which profile fields would move their relevance (an unverified Python skill). That is honest, actionable and does not leak other contractors' data or the model internals.
Two-sided acceptance and capacity
Ask: The owner picks from the list. Why does the contractor's preference, or their capacity, belong in the ranking at all?
Good answers name: Mutual score (P invite × P accept) plus capacity-aware re-rank with invite holds, Rank only on owner preference, Batch assignment (min-cost flow or stable matching) each hour, First come first served invites, no holds.
Our pick: Score = P(invite) × P(accept | invite) × E[success | hired], so a candidate who will decline sinks without being hidden. Each invite creates a 48 h hold in the capacity ledger; the re-ranker reads pending holds and open slots and demotes contractors whose pressure (holds per open slot) exceeds about 3, and drops anyone with zero open slots. Acceptance is a compare-and-set on the ledger (active < max) in one Lua script, so the last slot goes to exactly one project, and the loser hears "no longer available" at once. Declines carry a reason that updates the accept head's features the same day. Airbnb's published search-ranking work models host acceptance alongside guest booking, which is the same two-sided idea. If the platform ever auto-assigns (for example evaluator pools for AI training work), the same scores feed a min-cost flow solver with capacities, which is the batch version of this.
- A contractor accepts two invites within the same second and has one slot. What happens?
Both requests hit the accept script on the same ledger key. Redis runs scripts one at a time, so the first sees active < max and increments; the second sees active = max and returns 0. The second project gets an immediate 409 and the shortlist offers its next candidate. The engagement row is only written after the script succeeds, so the database cannot disagree. - The Redis ledger loses its data. What is the recovery?
The ledger is derived state: active counts come from engagements in Postgres and holds from pending invites. Rebuild it from a query, which is seconds for 2 M rows. During the rebuild, accepts fail closed (retry after a few seconds) rather than risking overbooking. Use Redis with AOF and a replica so this is rare. - Should an owner be told that a contractor has many pending invites?
Not the number, which leaks other clients' activity, but a hint like "responds within 1 day, currently in high demand" is useful and honest. Better still, the ranking has already lowered them, so the owner rarely needs the hint. - How do you measure whether the accept head is good?
Calibration on logged invites: of invites predicted at 0.3, about 30 % should be accepted, per cohort. And the business effect: the A/B metric invite-to-accept rate, which should rise when the head is on. Watch that it does not simply learn "cheap projects get declined" and bury low-budget openings; track fill rate by budget band.
Learning when only shown matches can happen
Ask: Your training data only contains people the old model chose to show. How do you learn anything the old model did not already believe?
Good answers name: Logged propensities, small exploration budget, inverse-propensity weighted training, Train on logged clicks and hires as they are, Position-bias correction only (click models), Full contextual bandit over every slot.
Our pick: Reserve 2 of the 20 shortlist positions (10 %) for exploration: candidates sampled from the eligible set outside the top 20, weighted toward those with little data (new, or few impressions) and never below a relevance floor, so an explored candidate is plausible rather than random. Every shown candidate is logged with its propensity: close to 1 for exploit slots, the sampling probability for explore slots. Training weights examples by 1 / propensity clipped at 20; offline evaluation uses self-normalised IPS on the exploration traffic to estimate hires under a candidate model. Labels are handled with their delays: an impression only becomes a negative for accept after its 48 h window closes, and for success after 30 days. Position bias inside the list is corrected with occasional randomised swaps of adjacent positions.
- How much does exploration cost, and how would you justify it to the business?
Measure it: exploration slots have lower invite rates, say half of exploit slots, so 10 % of slots costs ~5 % of invites on those lists. Against that, it is the only source of signal for new contractors and the only unbiased data for evaluating models. Show the counterfactual: models trained without it lose accuracy on new supply within months. Tune the share by segment: less in thin markets where each slot matters. - Propensities are tiny for some examples and the IPS estimate is all noise. What do you do?
Clip weights, use self-normalised IPS so the estimate is a weighted average rather than a sum, and report confidence intervals. Evaluate only on the exploration traffic, where propensities are bounded away from zero by design. If the interval is still too wide, the answer is more exploration traffic or a longer window, not a cleverer estimator. - A label arrives 30 days late. How do you train on recent data without it?
Separate heads with separate windows: invite and accept heads train on data a few days old; the success head trains on data older than 30 days. Recent impressions are excluded from the success head rather than treated as failures. Where needed, use an early proxy (still engaged at day 7) with a learned correction to the 30-day outcome. - Your offline estimate says the new model is +8 % hires. The A/B says +1 %. Which do you believe and why the gap?
The A/B, for the launch decision. Common causes of the gap: propensities logged wrong on some path (the fallback, say), interference because both arms share the same contractors, or the offline estimator evaluated on exploration traffic that is not representative of exploit traffic. Investigate each; the offline number is for ranking candidate models, not for predicting the size of the win.
Stopping popular contractors from monopolizing
Ask: Your best contractors win every ranking. Why is that a problem, and what do you do without hurting quality?
Good answers name: Capacity-proportional exposure: pressure penalty and daily budget in the re-ranker, Do nothing; let owners choose, Hard impression cap per contractor, Cohort exposure targets (e.g. new contractors, regions).
Our pick: In the re-ranker, multiply each contractor's score by a smooth pressure penalty once pending holds exceed about 3 per open slot, and apply a soft daily impression budget proportional to open slots, above which they only keep their position if their score lead over the next candidate is large. Zero open slots is a hard filter. Monitor daily: impression share of the top 1 %, Gini over impressions and over invites, invites per open slot, and accept rate by contractor decile. Launch changes to the penalty as opening-level A/B tests with hires per opening and 30-day success as the guardrails. The point to make in the interview: this is a marketplace-health control that lives after the model, so it can be tuned without retraining.
- An enterprise client insists on seeing their favourite contractor first every time. Now what?
Give them direct search and a "rehire" path that bypasses the recommendation list entirely; a known relationship is not a recommendation problem. The shortlist stays a discovery tool. If the contractor has capacity they accept; if not, the client learns that from the invite, which is the truth. - How do you tell whether spreading hurt quality?
The opening-level A/B: hires per opening, time to hire, and success at 30 days per arm. Also look at the contractors who gained exposure: their accept and success rates compared with the ones who lost it. If the second tier succeeds at nearly the same rate, spreading is almost free; if not, raise the pressure threshold. - At 10x scale, does concentration get better or worse?
Usually worse in absolute terms, because more openings compete for the same visible top tier and the feedback loop has more data to reinforce. The mechanisms scale fine, since they are per-contractor counters in the ledger, but the budget parameters need re-tuning per market, and cohort monitoring needs to be per skill cluster rather than global.
Cold start on both sides
Ask: A strong contractor joined yesterday with no history, and a new client posts their first project. How does either get a good match?
Good answers name: Content and vetting features plus exploration with a decaying boost, Rank newcomers with the main model and missing features as zero, A fixed "new" badge and permanent boost for 30 days.
Our pick: Make the model work without history: features include verified-skill assessment scores, vetting interview score, years in the skill, profile embedding similarity to the project, and a "history available" indicator so the model learns separate behaviour for newcomers instead of treating missing data as zero. Newcomers get priority inside the exploration slots, weighted by their predicted quality, so a well-vetted newcomer is explored first. The priority decays with impressions received (for example after 200 impressions or 3 invites they compete normally). For new owners, rank with project features and priors from similar owners (same company size, same skill set) until they have made a few invites.
- A newcomer gets 50 impressions and no invites. What does the system do?
Treat it as information: with propensities logged, 50 exploration impressions with no invites is a real signal and their score will fall through normal training. Before that, check whether it is fixable: an incomplete profile or unverified skills. Prompt them with specific actions. The boost ends by the decay rule either way; the system owes a fair look, not guaranteed work. - What stops someone from creating fresh accounts to keep getting the newcomer boost?
Vetting: verified identity and assessments before any exposure, so a new account costs real effort. Duplicate detection on identity documents, payment accounts and devices. And the boost is small and short relative to the effort of vetting. - How do you measure whether cold start is working?
Time to first invite and time to first hire for new contractors, by cohort and skill; share of impressions going to contractors under 30 days old; and the 90-day retention of new contractors. If time to first invite rises month over month, the exploration budget or the vetting features are failing.
Close · 5 min
Ask what breaks first at ten times the load, and what they would build next. Then give them your read: one thing that was strong, one thing that was missing, one thing to practise. Be specific; “good job” helps nobody.