Deployment strategies and safe releases
Shipping changes without outages: rolling deployments, blue-green, canary releases with automated analysis, feature flags versus deploys, backward-compatible changes across services and databases, progressive delivery across regions, rollbacks and deploy freezes.
Reading is half of it. See this used in a real interview: walk through Design a URL Shortener →
Most outages are caused by changes: a new release, a configuration push, a schema migration. Companies that deploy hundreds of times a day without constant incidents rely on deployment strategies that limit how many users a bad change can reach and how fast it can be undone. In interviews, a sentence on how you would roll out a risky change, and why it is safe, shows operational maturity.
Rolling deployment
Replace instances a few at a time: take some out of the load balancer, deploy, health-check, put them back, continue.
- No extra capacity needed beyond a small surge.
- Old and new versions run side by side during the rollout, so they must be compatible with each other.
- Rollback means another rolling deploy of the old version, which takes time.
The default for Kubernetes deployments.
Blue-green
Run two full environments: blue (current) and green (new). Deploy to green, test it, then switch traffic from blue to green at the load balancer or DNS.
- Instant rollback: switch back to blue.
- No mixed versions serving traffic at once (except shared dependencies like the database).
- Costs double capacity during the switch, and a cut-over sends all users to the new version at once.
Canary releases
Send a small share of traffic (1 %, then 5 %, 25 %, 100 %) to the new version, compare its metrics with the old version, and continue only if they are healthy.
- Automated canary analysis compares error rates, latency and business metrics between canary and baseline, with statistical checks, and rolls back automatically on regression.
- Limits the blast radius: a bad release affects 1 % of users for a few minutes.
- Needs good metrics and enough traffic for comparisons to mean something.
Netflix popularised automated canary analysis. See Design Netflix.
Deploy versus release
Deploying puts code in production; releasing exposes a feature to users. Feature flags separate them: deploy dark, then release gradually by flag, and turn it off without a deploy if it misbehaves. Deploys become routine and low-risk; risky changes are controlled by flags. See feature flags and A/B testing.
Compatibility rules
Because old and new versions coexist, and services deploy independently:
- APIs and events: additive changes only; new fields optional; consumers tolerate unknown fields. See schema evolution and serialization.
- Databases: expand and contract; never deploy code that requires a schema change in the same step that makes it. See online schema migrations.
- Clients: mobile apps stay old for months; the server must support them.
- Order of deploys: deploy consumers that understand a new field before producers that send it.
Progressive delivery across regions
Large systems roll out region by region, or by cell, with bake time between stages: first a small internal or canary region, then one production region, then the rest. A bad change is caught while it affects a fraction of users, and other regions remain healthy fallbacks. Configuration changes deserve the same treatment as code: many major cloud outages came from global config pushes. See multi-region architecture and multi-tenancy.
Rollbacks
- Make rollback fast and practised: one command or automatic.
- Keep changes reversible: data migrations and destructive changes are the hard part; decouple them from code deploys.
- Prefer roll back first, debug later when metrics degrade after a deploy.
- Sometimes roll forward (a quick fix) is safer, for instance after a one-way data change, but decide that deliberately.
Guardrails
- Automatic rollback on SLO burn during a rollout. See SLIs, SLOs and error budgets.
- Deploy freezes during peak events (holiday sales) or when the error budget is exhausted.
- Small, frequent deploys: easier to review, test and roll back than large batches.
- Every deploy and config change recorded, so incidents can be correlated with changes immediately.
Stateful and connection-heavy services
Services with long-lived connections (chat gateways) or in-memory state need gentle draining: stop new connections, migrate or let existing ones finish, and spread reconnects over time. See presence and connection management and Design WhatsApp.
In the interview
For a critical path like Design a Payment System or a high-traffic one like Design a URL Shortener: "changes ship behind flags and roll out as canaries with automated metric comparison, region by region, with automatic rollback; schema and API changes are backward compatible so old and new versions coexist."
Checklist
- Rolling deploys by default; blue-green where instant rollback is worth double capacity.
- Canaries with automated analysis and automatic rollback.
- Deploy separated from release with feature flags.
- Backward-compatible APIs, events and schemas; correct deploy order.
- Region-by-region or cell-by-cell rollout, config included.
- Fast, practised rollbacks; freezes during peaks.