SysDesignPrep.com
Study guide 175 of 183

Deployment strategies and safe releases

Shipping changes without outages: rolling deployments, blue-green, canary releases with automated analysis, feature flags versus deploys, backward-compatible changes across services and databases, progressive delivery across regions, rollbacks and deploy freezes.

Reading is half of it. See this used in a real interview: walk through Design a URL Shortener →

Most outages are caused by changes: a new release, a configuration push, a schema migration. Companies that deploy hundreds of times a day without constant incidents rely on deployment strategies that limit how many users a bad change can reach and how fast it can be undone. In interviews, a sentence on how you would roll out a risky change, and why it is safe, shows operational maturity.

Rolling deployment

Replace instances a few at a time: take some out of the load balancer, deploy, health-check, put them back, continue.

  • No extra capacity needed beyond a small surge.
  • Old and new versions run side by side during the rollout, so they must be compatible with each other.
  • Rollback means another rolling deploy of the old version, which takes time.

The default for Kubernetes deployments.

Blue-green

Run two full environments: blue (current) and green (new). Deploy to green, test it, then switch traffic from blue to green at the load balancer or DNS.

  • Instant rollback: switch back to blue.
  • No mixed versions serving traffic at once (except shared dependencies like the database).
  • Costs double capacity during the switch, and a cut-over sends all users to the new version at once.

Canary releases

Send a small share of traffic (1 %, then 5 %, 25 %, 100 %) to the new version, compare its metrics with the old version, and continue only if they are healthy.

  • Automated canary analysis compares error rates, latency and business metrics between canary and baseline, with statistical checks, and rolls back automatically on regression.
  • Limits the blast radius: a bad release affects 1 % of users for a few minutes.
  • Needs good metrics and enough traffic for comparisons to mean something.

Netflix popularised automated canary analysis. See Design Netflix.

Deploy versus release

Deploying puts code in production; releasing exposes a feature to users. Feature flags separate them: deploy dark, then release gradually by flag, and turn it off without a deploy if it misbehaves. Deploys become routine and low-risk; risky changes are controlled by flags. See feature flags and A/B testing.

Compatibility rules

Because old and new versions coexist, and services deploy independently:

  • APIs and events: additive changes only; new fields optional; consumers tolerate unknown fields. See schema evolution and serialization.
  • Databases: expand and contract; never deploy code that requires a schema change in the same step that makes it. See online schema migrations.
  • Clients: mobile apps stay old for months; the server must support them.
  • Order of deploys: deploy consumers that understand a new field before producers that send it.

Progressive delivery across regions

Large systems roll out region by region, or by cell, with bake time between stages: first a small internal or canary region, then one production region, then the rest. A bad change is caught while it affects a fraction of users, and other regions remain healthy fallbacks. Configuration changes deserve the same treatment as code: many major cloud outages came from global config pushes. See multi-region architecture and multi-tenancy.

Rollbacks

  • Make rollback fast and practised: one command or automatic.
  • Keep changes reversible: data migrations and destructive changes are the hard part; decouple them from code deploys.
  • Prefer roll back first, debug later when metrics degrade after a deploy.
  • Sometimes roll forward (a quick fix) is safer, for instance after a one-way data change, but decide that deliberately.

Guardrails

  • Automatic rollback on SLO burn during a rollout. See SLIs, SLOs and error budgets.
  • Deploy freezes during peak events (holiday sales) or when the error budget is exhausted.
  • Small, frequent deploys: easier to review, test and roll back than large batches.
  • Every deploy and config change recorded, so incidents can be correlated with changes immediately.

Stateful and connection-heavy services

Services with long-lived connections (chat gateways) or in-memory state need gentle draining: stop new connections, migrate or let existing ones finish, and spread reconnects over time. See presence and connection management and Design WhatsApp.

In the interview

For a critical path like Design a Payment System or a high-traffic one like Design a URL Shortener: "changes ship behind flags and roll out as canaries with automated metric comparison, region by region, with automatic rollback; schema and API changes are backward compatible so old and new versions coexist."

Checklist

  • Rolling deploys by default; blue-green where instant rollback is worth double capacity.
  • Canaries with automated analysis and automatic rollback.
  • Deploy separated from release with feature flags.
  • Backward-compatible APIs, events and schemas; correct deploy order.
  • Region-by-region or cell-by-cell rollout, config included.
  • Fast, practised rollbacks; freezes during peaks.

Open in your browser to sign in

Google does not allow sign-in inside this app's built-in browser. Open this page in Safari and sign in there. The link opens this same page.

Tap the ⋯ or share button at the top or bottom of the screen, then Open in browser. Or copy the link and paste it into Safari.