SysDesignPrep.com
Study guide 121 of 183

Migrating legacy systems with the strangler fig pattern

How to replace a monolith or legacy system without a big-bang rewrite: the strangler fig pattern, routing facades, extracting services by capability, data migration and synchronisation, dual running and shadow traffic, anti-corruption layers, and deciding what to migrate first.

Reading is half of it. See this used in a real interview: walk through Design a Payment System →

Few engineers design systems from scratch; most inherit one. "How would you migrate this monolith?" or "the current system cannot scale; how do you move to the new design?" are common senior and staff interview questions. Big-bang rewrites fail often: they take years, freeze features, and surprise everyone at cut-over. The strangler fig pattern (named after a vine that grows around a tree and gradually replaces it) replaces a system piece by piece while it keeps running.

The pattern

  1. Put a facade in front of the legacy system: a proxy, API gateway or routing layer that all traffic passes through. See API gateways.
  2. Build one capability in the new system.
  3. Route that capability's traffic to the new system, gradually.
  4. Repeat until the legacy system handles nothing, then retire it.

At every step, the product works, changes are small and reversible, and value is delivered early.

Choosing what to extract first

Good first candidates:

  • Painful and valuable: the part that limits scaling or slows every release.
  • Loosely coupled: few dependencies on the rest of the monolith's data and code.
  • Read-heavy or new features: easier to move than core write paths.

Avoid starting with the most entangled core (often orders or accounts); learn the process on something smaller. Map dependencies and data ownership first, often with domain modelling or event storming. See event-driven architecture and microservices.

Routing

The facade routes by path, operation, customer or percentage:

  • /search goes to the new search service; everything else to the monolith.
  • 1 % of users, then 10 %, then 100 %, with feature flags and instant rollback. See feature flags and A/B testing.
  • Internal callers inside the monolith are redirected too, by replacing in-process calls with calls to the new service behind an interface.

Data is the hard part

Code is easy to move; data is shared, large and constantly changing.

  • New service owns new data: for new capabilities, simple.
  • Moving existing data: copy it to the new store, then keep it in sync while both systems run, typically with change data capture from the legacy database.
  • Ownership flips once: at a planned point, the new service becomes the writer; the legacy system reads from it (or receives changes back via events) until it no longer needs the data.
  • Avoid long-term dual writes from application code: they drift. If you must, verify continuously.
  • Backfills and verification: batched copies and checksums or shadow reads before switching. See online schema migrations.

Proving the new system

  • Shadow traffic: send copies of real requests to the new system, discard its responses, and compare them with the legacy system's. Fix differences before switching.
  • Dual running with comparison: for critical calculations (pricing, billing), run both and alert on mismatches.
  • Gradual cut-over with clear success metrics and a rollback path.

See testing distributed systems and deployment strategies.

Anti-corruption layer

Legacy models are often awkward (overloaded tables, strange codes). Put a translation layer between the new service and the legacy system, so the new design is not shaped by old mistakes. It converts requests and data in both directions and disappears when the legacy part is retired.

Organisation and pace

  • Keep shipping features: build new features in the new system where possible, which naturally shifts weight over time.
  • Track progress by traffic and capabilities moved, not by lines of code.
  • Set a plan to actually delete legacy pieces; half-migrated systems are worse than either end state.
  • Expect the long tail (rare reports, admin tools, batch jobs) to take longer than the main paths.

Other migration patterns

  • Branch by abstraction: inside a codebase, introduce an interface, implement the new version behind it, switch, remove the old.
  • Parallel run: run old and new side by side for a period, compare outputs, then switch (common for financial systems).
  • Database-first migration: move the data store first (with CDC), then the code.

In the interview

"I would not rewrite it in one go. A routing facade goes in front; we extract the redirect path first, since it carries most traffic and has simple data; CDC keeps the new store in sync; shadow reads verify results; traffic shifts by percentage behind a flag; the old path is deleted once it has served nothing for a month. Then we move link creation and analytics the same way." This works for Design a URL Shortener scale-ups, payment platform rewrites (Design a Payment System) and monolith splits like Design DoorDash.

Checklist

  • A facade that can route any capability to old or new.
  • Extract loosely coupled, valuable pieces first.
  • Gradual, flag-controlled cut-over with rollback.
  • CDC-based data sync and a single ownership flip.
  • Shadow traffic and output comparison before switching.
  • Anti-corruption layer against legacy models.
  • Legacy pieces actually deleted.

Open in your browser to sign in

Google does not allow sign-in inside this app's built-in browser. Open this page in Safari and sign in there. The link opens this same page.

Tap the ⋯ or share button at the top or bottom of the screen, then Open in browser. Or copy the link and paste it into Safari.