SysDesignPrep.com
Study guide 85 of 183

Cell-based architecture

Limiting the blast radius of failures and bad deploys by splitting a system into independent cells: what a cell contains, routing customers to cells, sizing and placement, cell migration, shared global services, deployment by waves, and the costs of running many copies of a system.

Reading is half of it. See this used in a real interview: walk through Design Slack →

In a typical architecture, every customer shares every component, so one bad deploy, one poison request or one overloaded database can take down the whole product. A cell-based architecture splits the system into many independent, complete copies (cells), each serving a slice of customers. Failures stay inside one cell. AWS, Slack, Salesforce and many large platforms use cells. It is a strong staff-level answer to "how do you limit the blast radius?".

What a cell is

A cell is a full, self-contained deployment of the service: its own application servers, databases, caches and queues, sized for a fixed share of customers. Cells do not call each other in the request path. If cell 7 fails, customers in cells 1 to 6 and 8 to 40 are unaffected.

Compared with plain sharding (which splits only the database), cells split everything, so failures in any layer are contained.

Routing

A thin cell router in front of all cells maps each request to its cell:

  • The routing key is usually the tenant, workspace or account id (customers rarely need data from other customers). See multi-tenancy.
  • A cell directory stores the mapping, cached heavily at the router and at the edge.
  • The router must be extremely simple and robust, since it is shared by everyone: no business logic, minimal dependencies, static configuration where possible.

Sizing and placement

  • Choose a maximum cell size that you test at (for example, a few thousand tenants or a fixed request rate), and add cells rather than growing them. Known limits make capacity predictable.
  • Place cells across zones or regions as needed; a cell can itself span zones for availability.
  • Large customers may get dedicated cells; small ones share.

Cell migration

Moving a tenant between cells (to rebalance or to give a growing customer its own cell) requires copying their data and switching the directory entry, with techniques from online schema migrations: bulk copy, change data capture to catch up, a short write freeze or dual-writing, then flip and verify. Build this tooling early; it is essential for operating cells.

Shared global services

Some things cannot be split by tenant: authentication across tenants, billing, global search, the cell directory itself. Keep these few, simple and highly available, and design cells to keep working (perhaps degraded) when global services have problems. See rate limiting and resilience.

Deploying in waves

Cells enable safe deploys: release to one canary cell, then a small wave, then larger waves, with bake time and automatic rollback between waves. A bad release hurts a small fraction of customers. Configuration changes follow the same waves. See deployment strategies.

Shuffle sharding

A variation for shared resources: assign each customer a random combination of a few workers out of many. Two customers rarely share the same combination, so one misbehaving customer (a poison request, a traffic flood) affects only customers who share all of its workers, a tiny fraction. Useful for fleets like API front ends and queues.

Costs

  • Overhead: many copies of every component, each with headroom, cost more than one large shared system.
  • Operational complexity: more databases to patch, more dashboards, more deploy steps; heavy automation is required.
  • Cross-tenant features (global search, analytics across customers) need separate pipelines that aggregate from all cells. See data lakes and lakehouses.
  • Hot tenants still exist within a cell. See hot keys and skew.

Cells pay off at large scale, with many independent customers and high availability requirements. A small product does not need them.

In the interview

"The service is split into cells, each a complete stack serving up to a few thousand workspaces; a thin router maps workspace id to cell from a cached directory. Deploys roll out cell by cell with automatic rollback, so a bad change affects a few percent of customers. Big customers can have dedicated cells, and we have tooling to migrate workspaces between cells." See Design Slack and Design a Payment System.

Checklist

  • Complete, independent cells; no cross-cell calls in the request path.
  • A simple, robust router and cached cell directory keyed by tenant.
  • Fixed, tested maximum cell size; add cells to grow.
  • Tenant migration tooling between cells.
  • Few, simple global services; cells degrade gracefully without them.
  • Wave-based deploys and config changes; shuffle sharding for shared fleets.
  • Overhead and automation costs acknowledged.

Open in your browser to sign in

Google does not allow sign-in inside this app's built-in browser. Open this page in Safari and sign in there. The link opens this same page.

Tap the ⋯ or share button at the top or bottom of the screen, then Open in browser. Or copy the link and paste it into Safari.