Feature flags and A/B testing
Feature flags for safe releases, how flag evaluation works at scale, gradual rollouts and kill switches, A/B test design with consistent assignment, exposure logging, metrics and guardrails, network effects, and flag hygiene.
Reading is half of it. See this used in a real interview: walk through Design a News Feed →
Large products change dozens of times a day without breaking for everyone, because changes ship behind feature flags and are measured with experiments. Both come up in interviews as part of "how would you roll this out?" and "how would you know it worked?", and occasionally as a full design question ("design an experimentation platform").
Feature flags
A feature flag is a runtime switch that decides which code path runs, without a deploy:
- Release flags: ship code dark, then turn it on gradually.
- Kill switches: turn off a misbehaving feature in seconds.
- Operational flags: switch to a fallback, change a limit, shed a non-critical feature under load.
- Permission and entitlement flags: features for certain plans, tenants or beta users.
- Experiment flags: assign users to variants of an A/B test.
Evaluating flags at scale
Flags are checked on almost every request, so evaluation must be local and fast:
- A flag service stores rules (targeting by user, tenant, region, app version, percentage).
- SDKs in each service download the full rule set and evaluate locally in microseconds; they receive updates by streaming or polling (seconds).
- If the flag service is unreachable, SDKs keep the last known rules, with safe defaults for brand-new flags. A flag outage must never take down the product.
- Every change to a flag is audited (who, when, what), because flags are production changes.
Gradual rollouts
Roll out by percentage with consistent hashing of the user id (hash(flag, user) mod 100 < percentage), so each user stays in the same group as the percentage grows, and different flags are independent. A typical rollout: internal users, 1 %, 5 %, 25 %, 50 %, 100 %, watching error rates and latency at each step, with automatic rollback if guardrails break. See observability, operations and rollouts.
A/B testing
An experiment compares variants on randomly assigned users to measure the effect of a change.
Assignment: deterministic and sticky, by hashing the user (or device, for logged-out traffic) with the experiment id. Users must not flip between variants.
Exposure logging: record when a user actually saw the variant, not just that they were assigned, and analyse only exposed users. Otherwise effects get diluted by users who never reached the feature.
Metrics:
- A primary metric chosen before the test (conversion, retention, sessions per user).
- Guardrail metrics that must not get worse (latency, errors, unsubscribes, revenue).
- Per-metric statistics with confidence intervals, computed in the analytics warehouse from events joined to assignments. See OLTP versus OLAP.
Sample size and duration: decide them up front from the minimum effect worth detecting, run for whole weeks (behaviour differs by weekday), and do not stop early when the result looks good ("peeking" inflates false positives unless you use sequential testing methods).
When users are not independent
Standard A/B tests assume one user’s variant does not affect another’s outcome. That breaks in:
- Marketplaces: a new ranking for some riders changes which drivers are available to the others. See Design DoorDash.
- Social products: a feature for some users changes what their friends see.
- Shared resources: a change that uses more cache or capacity slows everyone.
Use switchback tests (alternate the whole region between variants by time window) or cluster randomisation (assign whole cities, or groups of connected users), at the cost of fewer independent units and more noise.
Running many experiments at once
Large platforms run hundreds of concurrent experiments. Layers keep experiments that could interfere (two ranking changes) mutually exclusive, while experiments in different layers (ranking and button colour) overlap freely. Holdout groups that see none of the quarter’s changes measure their combined effect.
Flag hygiene
Flags are technical debt: every flag doubles the paths through the code. Give every flag an owner and an expiry date, remove release flags once fully rolled out, and alert on flags that have been at 100 % for weeks. Stale flags have caused real outages when someone flipped one years later.
In the interview
For any risky change, say: behind a flag, rolled out by percentage with consistent assignment, guarded by error and latency metrics with automatic rollback, and (if it is a product change) measured by an A/B test with a primary metric and guardrails, using switchback or cluster randomisation if users interact.
Checklist
- Local flag evaluation with streamed updates and safe defaults.
- Audited flag changes and kill switches for risky features.
- Percentage rollouts with consistent hashing and automatic rollback.
- Sticky assignment, exposure logging, primary and guardrail metrics.
- Sample size decided up front; whole weeks; no peeking.
- Switchback or cluster designs where users interact.
- Owners and expiry dates for every flag.