Design a Project-to-Contractor Matching System
Match posted projects with qualified, available contractors across a 500 k-person talent pool, and learn from who actually gets hired.
Last updated 2026-09-26. Difficulty: hard. Patterns: recommendations, ranking, marketplace, exploration, search. Reported at Mercor, Upwork, LinkedIn, Indeed, Thumbtack.
The interviewer asks, the candidate answers and draws, and you press Next. Pause to answer yourself at the key decisions, and ask the AI Mentor anything along the way.
Functional requirements
- Project owners post openings with requirements. Skills, seniority, start date, hours per week, duration, timezone window, budget ceiling, and legal eligibility (work authorization, country, background check). Each requirement is marked must-have or nice-to-have.
- Contractors maintain a profile, availability and preferences. Verified skills, rate, weekly hours free, preferred project types, and the maximum number of concurrent engagements they will take.
- Owners get a ranked shortlist per opening. Top 20 candidates, each with reason codes ("4 of 5 required skills verified, 6 h timezone overlap"). This is a recommendation list, not automatic assignment: a human picks.
- Contractors get a feed of recommended projects. The same match, seen from the other side. A match only happens if both sides say yes, so the contractor feed is half of the product.
- Invite, accept and decline, with reasons. Owner invites, contractor accepts or declines within 48 h. Decline reasons ("rate too low", "not interested in the domain") are both a label and a preference update.
- Respect contractor capacity. A contractor at their concurrent-engagement limit must stop receiving invites, and one with several pending invites must not be shown to everyone.
- Learn from outcomes. Shown, invited, accepted, hired, and project success weeks later. The ranker improves from these without only reinforcing whoever it already showed.
- Out of scope. Payments and invoicing, the vetting interview itself, contracts, messaging, and fraud detection. Vetting results arrive as features.
Non-functional requirements
- Scale (2 M contractors · 500 k active · 15 k openings/day). Request rate is small. The difficulty is in what to show, not how fast.
- Shortlist latency (p99 < 800 ms). An owner waits on a page load. Retrieval, scoring 1 k candidates and re-ranking must fit together.
- Eligibility correctness (0 ineligible candidates shown). Showing someone without work authorization for the country, or already at capacity, is a real harm and a legal problem. This is the requirement that drives the hardest tradeoff: hard rules must never be traded off against a learned score.
- Freshness (availability change visible < 1 min). A contractor who just accepted their second project and hit their limit must drop out of shortlists quickly, or owners invite someone who cannot say yes.
- Match quality (invite → accept > 40 % · hire → success > 80 %). Measured per cohort. Accept rate is the two-sided signal; success is the one that matters but arrives weeks late.
- Fair exposure (new contractors get ≥ 5 % of impressions). A marketplace that only shows proven people starves its supply. Concentration is monitored like latency.
- Explainability and audit (every impression reproducible). Both sides need a reason for a recommendation, and an audit must be able to replay why a person was or was not shown. Log model version, features and score with each impression.
Back-of-envelope estimates
- Ranking requests: ~14/s avg · ~55/s peak. 500 k active contractors × 2 feed loads a day = 1 M, plus 200 k owner shortlist loads = 1.2 M/day ÷ 86 400 ≈ 14/s, ×4 at peak. Throughput is not the hard part.
- Candidates after hard filters: ~2 k per opening. 500 k active × ~5 % with the required skills ≈ 25 k; availability, timezone, rate and eligibility keep ~1 in 12 ≈ 2 k. Retrieval hands the top 1 k to the ranker.
- Scoring work: ~55 k scores/s peak. 55 requests/s × 1 k candidates = 55 k model evaluations/s. A gradient-boosted model scores ~100 k rows/s per core, so a few cores. Feature fetches (1 k keys per request) dominate the latency, not the model.
- Embedding index: ~2 GB. 2 M contractors × 256-dim float32 vector = 2 M × 1 KB = 2 GB. One in-memory ANN index per replica, rebuilt nightly and patched on profile change.
- Interaction events: ~25 M/day · ~5 TB/yr. 1.2 M ranking requests × 20 shown = 24 M impressions/day, plus clicks, invites and outcomes ≈ 25 M events × ~500 B (with score, position, propensity, model version) ≈ 12 GB/day ≈ ~5 TB/yr.
- Concurrent engagements: ~450 k. 15 k openings/day × ~30-day average engagement ≈ 450 k live engagements. With a typical limit of 2 each, that needs at least 225 k distinct contractors, 45 % of the active pool. Capacity binds: the same top 5 % cannot fill the market.
- Labels per month: ~450 k hires · success at +30 days. 15 k hires/day × 30 = 450 k hire labels a month against ~720 M impressions: a positive rate of ~0.06 %. Success labels arrive a month later. So the ranker also trains on nearer signals (invite, accept) and weights them.
Components
- Owner app: Where a project owner posts openings, reads the ranked shortlist with reason codes, relaxes constraints when the list is thin, and sends invites.
- Contractor app: Profile, availability and preferences, a feed of recommended projects, and the inbox where invites are accepted or declined with a reason. The contractor side is half the match: a list the contractor would decline is a bad list.
- Email / push (SES · APNs · FCM): Delivers invites and reminders. Invites expire after 48 h, so delivery latency and the reminder at 24 h directly move the accept rate.
- Profile service (skill taxonomy): Owns contractor profiles and project specs. Normalises free-text skills onto a taxonomy ("ReactJS", "React.js" → react) and records which skills are verified by assessment. Separates must-have from nice-to-have on every project requirement at write time, so the matcher never has to guess.
- API gateway (auth · rate limits): Authenticates owners and contractors, routes shortlist and feed requests to the match service and invite actions to the allocation service. Stamps a request id that follows the impression into the logs.
- Core DB (Postgres): Source of truth for contractors, projects, openings, invites, engagements and outcomes. Change data capture streams every row change onto the event log, which feeds both the search index and the warehouse.
- Match service (stateless orchestrator): Runs one request end to end: compile the opening into hard filters, retrieve, score, re-rank for capacity and exposure, attach reason codes, and log the impression with its propensity. Hard constraints are applied here as filters and are never inputs the model can outweigh.
- Exposure & capacity (re-ranker · invites): Takes the ranked list and makes it safe for the marketplace: removes anyone at capacity, demotes contractors with many pending invites, caps daily exposure, and reserves exploration slots. Also owns the invite lifecycle and the capacity holds that go with it.
- Event log (Kafka): Impressions (with position, score, propensity and model version), clicks, invites, responses and CDC changes from the core DB. Keyed by contractor id so a contractor's availability changes apply in order.
- Candidate index (OpenSearch · filters + kNN): One document per contractor: taxonomy skills, availability, timezone, rate, eligibility flags, open capacity and a profile embedding. Serves filtered retrieval: hard constraints as filter clauses, then lexical and vector similarity inside the survivors. Updated from CDC within seconds.
- Ranker (GBDT, multi-task): Scores ~1 k candidates with separate heads for P(owner invites), P(contractor accepts) and P(success | hired), combined into one expected-value score. Returns the per-feature contributions that become reason codes.
- Capacity ledger (Redis · slots + holds): Per contractor: max concurrent engagements, active count, pending invite holds with TTL, and impressions today. Every accept is a compare-and-set here so two projects cannot both take the last slot.
- Warehouse (S3 + Iceberg): Every impression and outcome, joined by impression id, with delayed labels filled in as they arrive (accept in days, success in a month). The training set and the audit trail are the same tables.
- Feature store (online Redis · offline Iceberg): Contractor features (verified skills, response rate, past success, rate percentile), project features and pair features (skill overlap, timezone overlap). Online and offline values come from the same definitions, so training and serving see the same numbers.
- Training & eval (daily · counterfactual): Retrains the ranker from logged impressions weighted by inverse propensity, evaluates new models counterfactually on exploration traffic before any A/B test, and runs the concentration and cohort-exposure monitors.
User flows
- Owner opens a ranked shortlist. Filter, retrieve, score, re-rank, explain, log. The order of those steps is the design: eligibility comes first and is never part of the learned score.
- Owner opens the shortlist for an opening. The request carries the opening id only. Everything the match needs is read server-side, so a client cannot widen its own filters or see contractors it should not.
- Match service loads the opening and compiles its must-haves into filters. Must-haves become boolean filter clauses: verified skills ⊇ required, available hours ≥ requested, timezone overlap ≥ 4 h, rate ≤ budget ceiling, eligibility flags for the project's country, open capacity > 0. Nice-to-haves become ranking features instead. The split is decided by the owner at posting time, not inferred by the model.
- Candidate index returns the top 1 k eligible contractors by relevance. One query: filters first, which cut 500 k active to ~2 k, then hybrid relevance inside the survivors (BM25 on skills and titles plus kNN on the profile embedding). Filtering before the vector search matters: a vector search that filters afterwards can return 1 k near neighbours of which 5 are eligible.
- Ranker scores the 1 k with features and combines its heads into one expected value. Score = P(invite) × P(accept | invite) × E[success | hired]. The accept head is what makes this two-sided: a candidate the owner loves who will decline because the rate is 30 % below their floor is a wasted slot. One batched feature read of 1 k contractor keys plus the project key, ~15 ms.
- Exposure & capacity re-ranks the list for the marketplace. Drops anyone whose slots filled since indexing, demotes contractors whose pending invites already exceed their open slots, applies a daily exposure budget, and fills 2 of 20 positions with exploration candidates sampled from lower-ranked eligible contractors. Each position records the probability it was shown there: that propensity is what makes the logs learnable.
- Match service attaches reason codes and returns the top 20. Reasons come from the must-have filter (all satisfied, stated plainly) and the top positive feature contributions ("shipped 3 RLHF projects", "8 h overlap with your team"). Never a reason derived from a protected attribute, and never "because others clicked".
- The impression is logged with positions, scores, propensities and model version. Also the ids that passed the filters but were not shown, sampled. Without the not-shown set and the propensities, the training data can only say who was chosen from among the people the old model liked.
- Invite, accept, and the capacity race. A match needs both sides to agree, and a contractor has a finite number of slots. Holds and a compare-and-set keep the capacity promise honest.
- Owner invites three contractors from the shortlist. Each invite places a soft hold on the contractor with a 48 h TTL. Holds do not block acceptance elsewhere, but they count toward "pending pressure", which the re-ranker uses to stop showing a contractor who already has five invites and one open slot.
- The invite is stored and the contractor is notified. The invite row carries the impression id and the reasons shown to the owner, so the contractor sees why they were picked. A reminder fires at 24 h; unanswered invites expire at 48 h and count as a soft decline.
- Contractor accepts.
- Allocation commits the engagement only if a slot is still free. Compare-and-set on the ledger: active < max, then increment and release this invite's hold, in one script. Then the engagement row is written. If two projects race for the last slot, exactly one wins and the other gets a clear "no longer available" instead of a silent overbooking.
- The capacity change reaches the index, so the contractor stops appearing where they cannot say yes. CDC turns the engagement insert into an index update in seconds. The ledger check at re-rank time covers the gap, so the index being a few seconds stale shows up as a slightly shorter list, never as an ineligible candidate.
- Another contractor declines with a reason; the reason is logged as a label and a preference. "Rate too low" updates the contractor's effective rate floor; "not my domain" lowers similar projects in their feed; "no time" should have been caught by availability, so it prompts them to update it. Declines are the cheapest honest signal the system gets.
- A contractor who declined for lack of time updates their availability. The decline screen links straight to the availability form, prefilled. Accurate availability is worth more than any model feature: it removes a whole class of invites that could never be accepted. The change reaches the index through the same CDC path as an engagement.
- A star contractor floods every shortlist. The best-scored profiles win every ranking, collect dozens of invites they cannot accept, and starve everyone else. The fix lives in the re-ranker, not the model.
- Monitoring shows the top 1 % of contractors receiving 35 % of impressions and invites. Concentration is tracked daily: impression share of the top 1 %, a Gini coefficient over impressions, and median invites per contractor per open slot. The symptom that owners feel is a falling accept rate, because the people they invite are drowning in invites.
- The re-ranker compares each contractor's pending invites with their open slots. Pressure = pending holds ÷ open slots. Above 3, the contractor's score is multiplied down smoothly (for example by 1 / (1 + pressure - 3)). They are not hidden, just no longer shown first to everyone.
- A daily exposure budget spreads impressions in proportion to capacity. Budget ≈ k × open slots per day. Past it, the contractor can still appear when they are far ahead of the next candidate, but ties go to someone with room. This turns a winner-take-all ranking into one that respects the fact that one person can only say yes twice.
- When the star's slots fill, the capacity change removes them from retrieval entirely. open_slots = 0 is a hard filter, so a fully booked contractor disappears from all shortlists within seconds. Their feed still works; they can see projects starting after their current engagement ends if they set a future availability date.
- The quality cost of spreading is measured, not assumed. An A/B test at the opening level compares accept rate, time to hire and 30-day success between the capped and uncapped re-ranker. Typical outcome on two-sided marketplaces: accept rate rises because invites go to people who can say yes, and success barely moves because the second-best candidate is usually nearly as good.
- No one meets every hard constraint. The honest answer is an empty list with a map of which constraint to relax, not a list that quietly ignores a must-have. Plus the ranker being down.
- Retrieval with every must-have returns 3 candidates. Too few to be useful. The match service never silently drops a must-have to fill the list: the owner said it was non-negotiable, and a candidate without US work authorization is not a near miss.
- Match service runs count queries with each negotiable must-have relaxed in turn. Only the constraints the owner may relax: budget, hours, timezone, a secondary skill. Legal eligibility and certification requirements are never candidates for relaxation. Counts are cheap (no scoring) and run in parallel, ~6 queries.
- Owner sees the 3 matches plus the tradeoffs that would widen the list. "3 match everything. Raising the budget to $150/h adds 48; accepting 2 h of timezone overlap adds 21." The owner decides. This converts a dead end into a product conversation and gives the marketplace a demand signal about what supply is missing.
- The thin result is logged as unmet demand for supply planning. Openings with fewer than 10 eligible candidates, grouped by skill and region, tell the supply team whom to recruit. A matching system that only ranks never learns what it cannot match.
- Separately, the ranker times out on a request. Budget 250 ms. On timeout the match service falls back to retrieval order with a simple rule score (verified skill overlap, then response rate), still after the hard filters and the capacity re-rank. Eligibility never depends on the model being up.
- The fallback impression is logged with its own policy name. model: "fallback-rules" with its own propensities. Training either uses these impressions with the right weights or excludes them; mixing them in as if the main model had chosen them corrupts the counterfactual estimates.
- Learning from outcomes without exposure bias. Only shown contractors can be hired, so naive training teaches the model to prefer whoever it already showed. Propensities, exploration and delayed labels fix that.
- Impressions, responses and engagement outcomes land in the warehouse. Impressions from the match service, invites and responses and engagement completions from CDC. Everything joins on impression_id and contractor_id, so every label points back to the exact list it came from.
- Delayed labels fill in as they arrive. Invite within the session, accept within 48 h, hire within ~a week, success at 30 days (completed, owner rating ≥ 4, or extended). Training on success alone would wait a month and see 450 k positives; so each head trains on its own label, and recent impressions are held back until their label window closes rather than counted as negatives.
- Training weights each shown example by the inverse of its propensity. An exploration candidate shown with probability 0.02 who gets hired counts 50 times as much as a top-ranked one shown with probability 1. Weights are clipped (e.g. at 20) to control variance. This corrects for "the model only learned about people it chose to show".
- A candidate model is evaluated offline on the exploration traffic. Self-normalised inverse propensity scoring estimates how many accepts and hires the new ranking would have produced. Offline NDCG on the logged lists is reported too but not trusted alone, because it rewards reproducing the old model's choices.
- Cohort and concentration checks gate the release. Exposure share for new contractors, by region and by rate band, and top-1 % concentration, compared with the current model. A model that wins on accept rate by showing only veterans fails the gate.
- The model ships to the ranker in shadow, then to an opening-level A/B. Shadow mode scores live traffic without serving, checking latency and score drift. The A/B randomises by opening, not by user, because two owners competing for the same contractor interfere: a test split by user leaks through shared supply. Primary metric: hires per opening at 14 days; guardrails: accept rate, exposure share, success at 30 days.
Deep dives
Keep eligibility out of the learned score
Why not just give the model timezone, rate and work authorization as features and let it learn what matters?
A learned score is a weighted compromise: it will happily rank a contractor with 1 h of timezone overlap first if everything else about them is excellent, because that is what "learn what matters" means. For a nice-to-have that is exactly right. For a must-have it is a bug. Showing someone without work authorization for the country, or someone already at their engagement limit, is not a slightly worse match; it is a match that cannot happen, and in the authorization case a legal problem.
The key observation is that constraints come in two kinds that look alike in the data: non-negotiable (legal eligibility, certification, capacity, and whatever the owner explicitly marked must-have) and preferences (seniority, nice-to-have skills, rate within budget). The system has to know which is which before ranking, and the owner is the one who knows for their project.
- Hard filters before retrieval; preferences as ranking features chosen
- Everything as features in one learned score rejected
- Learned score with large penalty features for violations rejected
- Post-ranking filter that drops violators from the top 20 situational: as a final check on top of retrieval filters, for fast-changing state like capacity
The answer: Classify every requirement at write time: legal eligibility and certification are always hard, capacity is always hard, and the owner marks the rest as must-have or nice-to-have in the posting form. Must-haves compile to filter clauses in the retrieval query so the ranker never sees an ineligible candidate; nice-to-haves become features. Fast-changing hard state (open slots) is checked again at re-rank time against the capacity ledger because the index can lag by seconds. When the filters leave too few candidates, the system reports which negotiable must-have to relax and how many people that adds, and never relaxes one on its own. LinkedIn's published Recruiter search work follows the same split: structured filters first, learned relevance inside them.
An owner marks 9 of 10 requirements as must-have and gets zero candidates. What now?
Show zero honestly, then the relaxation map: which single constraint, relaxed by how much, adds how many eligible people. Most owners pick one. Over time, track which must-haves owners relax most often; the posting form can then pre-suggest them as nice-to-have. The system never overrules the owner, but it can make the cost of strictness visible.
A contractor's availability in the index says 30 h/week but they are actually fully booked elsewhere off-platform. How do you catch it?
You cannot see off-platform work, so rely on the contractor: ask for availability confirmation every 2 weeks and decay "available" to unknown if they do not respond. Declines with reason "no capacity" flip it immediately. And treat stale availability as a ranking feature (a confidence), not the hard filter itself: the hard filter is "not known to be unavailable".
Which constraints would you let the model learn to soften, if any?
Only ones the owner did not mark: a rate slightly over the soft target, a secondary skill, seniority. The model can learn that owners frequently hire a mid-level engineer when they asked for senior, and rank accordingly among eligible candidates. It can never move a candidate across a must-have boundary.
How do you prove to an auditor that nobody ineligible was shown?
Each impression log carries the compiled filter and the snapshot version of each shown contractor's eligibility fields. A daily job re-evaluates the filter against the logged values and alerts on any violation. Because the filter is plain code, the auditor can read it; you could not make that argument about a model weight.
Retrieval then ranking
You have 500 k active contractors. Why not just score all of them with your best model?
Scoring 500 k candidates with feature fetches per request is ~500 k feature reads and model evaluations for a page load: tens of seconds, and almost all of it wasted on people who fail a hard filter or are obviously irrelevant. The standard answer is a funnel: cheap, high-recall retrieval narrows to ~1 k, and an expensive, high-precision ranker orders those.
The retrieval stage must not miss good people, which is harder than it sounds with skills. "PyTorch" and "deep learning research" are the same person in two vocabularies. Lexical matching on a taxonomy catches exact skills; an embedding catches the semantic neighbours. The ranker then decides what the platform actually wants, which in a two-sided market is not "who is most relevant" but "which match will both sides accept and succeed at".
- Filtered hybrid retrieval (taxonomy + embedding) then a multi-task ranker chosen
- Score every active contractor per request rejected
- Hand-weighted rule score (skill overlap, rating, response rate) situational: at launch before there are enough outcomes, and as the fallback when the ranker is down
- LLM re-ranking of the top 50 against the project description situational: as an extra feature or explanation generator on a small top-k, not as the primary ranker
The answer: Retrieval is one OpenSearch query: must-have filter clauses, then BM25 on taxonomy skills and titles plus kNN on a 256-dim profile embedding over the filtered set, merged by reciprocal rank fusion, top 1 k. Recall is measured offline weekly: of the contractors who were hired in the last month, what fraction would the retrieval stage have returned for that opening; target above 95 %. The ranker is a gradient-boosted multi-task model with three heads (P(invite), P(accept | invite), E[success | hired]) whose product is the score; GBDT because the features are tabular, it trains in minutes, and it gives per-feature contributions for reason codes. An embedding two-tower model is a natural next step once there is enough data, used as a retrieval source rather than a replacement for the ranker.
How would you know retrieval is dropping good candidates?
Retrieval recall at 1 k against a proxy of "good": contractors who were later hired for similar openings, or who were invited from exploration slots and accepted. If recall drops below target for a skill cluster, the usual cause is taxonomy drift (a new framework name) or an embedding that has not seen the new vocabulary. Also watch the share of hires that came from outside the main retrieval path, such as direct search.
Why multiply the heads instead of training one model on "hired"?
Hired is a rare, late label that mixes two decisions: the owner's and the contractor's. Separate heads train on denser labels (invite and accept happen within days), each is easier to diagnose, and the product lets you change the objective (weight success more) without retraining. The cost is that errors compound, so calibration of each head matters.
The ranker is 10x slower after adding pair features. What do you do?
First cut the candidate count: score 1 k with a light first-stage model, then 200 with the heavy one. Precompute the pair features that do not depend on the request (contractor x owner history) nightly. And batch the feature store reads, which are usually the real cost, not the trees.
A contractor asks why they never appear for Python projects. What can you tell them?
Which hard filters they fail most often across recent Python openings (for example "hours available below requested" in 60 % of them), and which profile fields would move their relevance (an unverified Python skill). That is honest, actionable and does not leak other contractors' data or the model internals.
Two-sided acceptance and capacity
The owner picks from the list. Why does the contractor's preference, or their capacity, belong in the ranking at all?
A match only exists if both sides say yes, and a contractor can hold only a few engagements. If the ranking only models the owner, it fills shortlists with people who will decline (rate too low, wrong domain) or who already have ten invites for two slots. Each wasted invite costs the owner two days and lowers the accept rate for everyone.
At our numbers capacity really binds: 450 k concurrent engagements need at least 225 k distinct contractors, nearly half the active pool. So the system is an allocation problem wearing a recommendation costume. The question is how much of the allocation to do centrally without taking the choice away from the humans.
- Mutual score (P invite × P accept) plus capacity-aware re-rank with invite holds chosen
- Rank only on owner preference rejected
- Batch assignment (min-cost flow or stable matching) each hour situational: when the platform assigns work itself, for example high-volume task work or cohort placements
- First come first served invites, no holds rejected
The answer: Score = P(invite) × P(accept | invite) × E[success | hired], so a candidate who will decline sinks without being hidden. Each invite creates a 48 h hold in the capacity ledger; the re-ranker reads pending holds and open slots and demotes contractors whose pressure (holds per open slot) exceeds about 3, and drops anyone with zero open slots. Acceptance is a compare-and-set on the ledger (active < max) in one Lua script, so the last slot goes to exactly one project, and the loser hears "no longer available" at once. Declines carry a reason that updates the accept head's features the same day. Airbnb's published search-ranking work models host acceptance alongside guest booking, which is the same two-sided idea. If the platform ever auto-assigns (for example evaluator pools for AI training work), the same scores feed a min-cost flow solver with capacities, which is the batch version of this.
A contractor accepts two invites within the same second and has one slot. What happens?
Both requests hit the accept script on the same ledger key. Redis runs scripts one at a time, so the first sees active < max and increments; the second sees active = max and returns 0. The second project gets an immediate 409 and the shortlist offers its next candidate. The engagement row is only written after the script succeeds, so the database cannot disagree.
The Redis ledger loses its data. What is the recovery?
The ledger is derived state: active counts come from engagements in Postgres and holds from pending invites. Rebuild it from a query, which is seconds for 2 M rows. During the rebuild, accepts fail closed (retry after a few seconds) rather than risking overbooking. Use Redis with AOF and a replica so this is rare.
Should an owner be told that a contractor has many pending invites?
Not the number, which leaks other clients' activity, but a hint like "responds within 1 day, currently in high demand" is useful and honest. Better still, the ranking has already lowered them, so the owner rarely needs the hint.
How do you measure whether the accept head is good?
Calibration on logged invites: of invites predicted at 0.3, about 30 % should be accepted, per cohort. And the business effect: the A/B metric invite-to-accept rate, which should rise when the head is on. Watch that it does not simply learn "cheap projects get declined" and bury low-budget openings; track fill rate by budget band.
Learning when only shown matches can happen
Your training data only contains people the old model chose to show. How do you learn anything the old model did not already believe?
Every label in this system is conditional on an impression. A contractor ranked 400th was never seen, so never invited, so appears in the data as a non-event rather than a negative. Train on that naively and the new model learns to agree with the old one: a feedback loop in which proven contractors get more exposure, more hires and higher scores, and everyone else sinks, regardless of how good they are.
The fix needs two things the logs usually lack: the probability with which each shown candidate was shown (so each outcome can be reweighted to what a uniform policy would have seen), and some deliberate randomness (so there are outcomes at all for people the model does not favour).
- Logged propensities, small exploration budget, inverse-propensity weighted training chosen
- Train on logged clicks and hires as they are rejected
- Position-bias correction only (click models) situational: together with exploration, for the within-list bias
- Full contextual bandit over every slot situational: for the exploration slots specifically, choosing whom to explore
The answer: Reserve 2 of the 20 shortlist positions (10 %) for exploration: candidates sampled from the eligible set outside the top 20, weighted toward those with little data (new, or few impressions) and never below a relevance floor, so an explored candidate is plausible rather than random. Every shown candidate is logged with its propensity: close to 1 for exploit slots, the sampling probability for explore slots. Training weights examples by 1 / propensity clipped at 20; offline evaluation uses self-normalised IPS on the exploration traffic to estimate hires under a candidate model. Labels are handled with their delays: an impression only becomes a negative for accept after its 48 h window closes, and for success after 30 days. Position bias inside the list is corrected with occasional randomised swaps of adjacent positions.
How much does exploration cost, and how would you justify it to the business?
Measure it: exploration slots have lower invite rates, say half of exploit slots, so 10 % of slots costs ~5 % of invites on those lists. Against that, it is the only source of signal for new contractors and the only unbiased data for evaluating models. Show the counterfactual: models trained without it lose accuracy on new supply within months. Tune the share by segment: less in thin markets where each slot matters.
Propensities are tiny for some examples and the IPS estimate is all noise. What do you do?
Clip weights, use self-normalised IPS so the estimate is a weighted average rather than a sum, and report confidence intervals. Evaluate only on the exploration traffic, where propensities are bounded away from zero by design. If the interval is still too wide, the answer is more exploration traffic or a longer window, not a cleverer estimator.
A label arrives 30 days late. How do you train on recent data without it?
Separate heads with separate windows: invite and accept heads train on data a few days old; the success head trains on data older than 30 days. Recent impressions are excluded from the success head rather than treated as failures. Where needed, use an early proxy (still engaged at day 7) with a learned correction to the 30-day outcome.
Your offline estimate says the new model is +8 % hires. The A/B says +1 %. Which do you believe and why the gap?
The A/B, for the launch decision. Common causes of the gap: propensities logged wrong on some path (the fallback, say), interference because both arms share the same contractors, or the offline estimator evaluated on exploration traffic that is not representative of exploit traffic. Investigate each; the offline number is for ranking candidate models, not for predicting the size of the win.
Stopping popular contractors from monopolizing
Your best contractors win every ranking. Why is that a problem, and what do you do without hurting quality?
Ranking is winner-take-all by construction: if two contractors score 0.91 and 0.90, one of them gets position 1 in every shortlist. Across thousands of openings that turns a tiny quality difference into a huge exposure difference. The top contractors get dozens of invites for two slots and decline most, owners see a lower accept rate, and the next-best contractors get nothing, lose interest and leave. The supply side erodes.
The key observation is that exposure is only valuable where it can convert. A contractor with zero open slots gains nothing from being shown, and one with seven pending invites for one slot gains almost nothing. So capacity, not a fairness quota, is the natural unit for spreading exposure.
- Capacity-proportional exposure: pressure penalty and daily budget in the re-ranker chosen
- Do nothing; let owners choose rejected
- Hard impression cap per contractor situational: as a safety net above the soft budget, against runaway feedback
- Cohort exposure targets (e.g. new contractors, regions) situational: for new-contractor ramp-up, implemented through the exploration slots rather than as quotas on the main list
The answer: In the re-ranker, multiply each contractor's score by a smooth pressure penalty once pending holds exceed about 3 per open slot, and apply a soft daily impression budget proportional to open slots, above which they only keep their position if their score lead over the next candidate is large. Zero open slots is a hard filter. Monitor daily: impression share of the top 1 %, Gini over impressions and over invites, invites per open slot, and accept rate by contractor decile. Launch changes to the penalty as opening-level A/B tests with hires per opening and 30-day success as the guardrails. The point to make in the interview: this is a marketplace-health control that lives after the model, so it can be tuned without retraining.
An enterprise client insists on seeing their favourite contractor first every time. Now what?
Give them direct search and a "rehire" path that bypasses the recommendation list entirely; a known relationship is not a recommendation problem. The shortlist stays a discovery tool. If the contractor has capacity they accept; if not, the client learns that from the invite, which is the truth.
How do you tell whether spreading hurt quality?
The opening-level A/B: hires per opening, time to hire, and success at 30 days per arm. Also look at the contractors who gained exposure: their accept and success rates compared with the ones who lost it. If the second tier succeeds at nearly the same rate, spreading is almost free; if not, raise the pressure threshold.
At 10x scale, does concentration get better or worse?
Usually worse in absolute terms, because more openings compete for the same visible top tier and the feedback loop has more data to reinforce. The mechanisms scale fine, since they are per-contractor counters in the ledger, but the budget parameters need re-tuning per market, and cohort monitoring needs to be per skill cluster rather than global.
Cold start on both sides
A strong contractor joined yesterday with no history, and a new client posts their first project. How does either get a good match?
Behavioural features (response rate, past success, owner ratings) are the strongest predictors, and a newcomer has none. A ranker that leans on them buries every new contractor under veterans with similar skills, and the newcomer, seeing no invites in their first week, leaves. New owners have the mirror problem: no history of whom they invite or hire.
The newcomer is not unknown, only unobserved. Vetting results, verified skills and the profile text say a lot. The design question is how to use what is known now and how to earn behavioural data quickly and safely.
- Content and vetting features plus exploration with a decaying boost chosen
- Rank newcomers with the main model and missing features as zero rejected
- A fixed "new" badge and permanent boost for 30 days situational: at launch, before exploration and vetting features exist
The answer: Make the model work without history: features include verified-skill assessment scores, vetting interview score, years in the skill, profile embedding similarity to the project, and a "history available" indicator so the model learns separate behaviour for newcomers instead of treating missing data as zero. Newcomers get priority inside the exploration slots, weighted by their predicted quality, so a well-vetted newcomer is explored first. The priority decays with impressions received (for example after 200 impressions or 3 invites they compete normally). For new owners, rank with project features and priors from similar owners (same company size, same skill set) until they have made a few invites.
A newcomer gets 50 impressions and no invites. What does the system do?
Treat it as information: with propensities logged, 50 exploration impressions with no invites is a real signal and their score will fall through normal training. Before that, check whether it is fixable: an incomplete profile or unverified skills. Prompt them with specific actions. The boost ends by the decay rule either way; the system owes a fair look, not guaranteed work.
What stops someone from creating fresh accounts to keep getting the newcomer boost?
Vetting: verified identity and assessments before any exposure, so a new account costs real effort. Duplicate detection on identity documents, payment accounts and devices. And the boost is small and short relative to the effort of vetting.
How do you measure whether cold start is working?
Time to first invite and time to first hire for new contractors, by cohort and skill; share of impressions going to contractors under 30 days old; and the 90-day retention of new contractors. If time to first invite rises month over month, the exploration budget or the vetting features are failing.