Event sourcing, CQRS and change data capture
Storing events instead of state, separating write and read models (CQRS), keeping read models in sync with change data capture and the transactional outbox, and when these patterns are worth their complexity.
Reading is half of it. See this used in a real interview: walk through Design a Payment System →
Three patterns travel together and are often confused. Event sourcing stores what happened instead of the current state. CQRS uses different models for writing and for reading. Change data capture turns a normal database’s changes into a stream. You will use CDC and simple CQRS in many designs; full event sourcing in a few. Knowing which is which, and what each costs, is what interviewers look for.
Event sourcing: the log is the truth
In a normal design, an account row holds balance = 120. In an event-sourced design, the store holds the facts: Deposited 100, Deposited 50, Withdrew 30. The current balance is computed by replaying them.
What you gain:
- A complete, immutable audit trail: you know how every value came to be. Banks, ledgers and trading systems work this way by nature.
- Time travel: rebuild the state as of any moment, which helps debugging and regulators.
- New read models from old data: add a projection next year and replay history into it.
- Natural fit for streams: the events are already what other systems want to consume.
What it costs:
- Reads need projections: you rarely replay on every request, so you maintain derived views (below).
- Long histories: replaying millions of events is slow, so you keep snapshots (state as of event N) and replay only what follows.
- Schema evolution: events are forever, so old event versions must be readable or upcast when the format changes.
- Deletes are hard: "forget this user" conflicts with an immutable log; you need crypto-shredding (encrypt per user, delete the key) or compaction strategies.
- Unfamiliar: most teams and tools assume current-state tables.
Use it where history is the product (ledgers, order books, document edit histories) or required by audit. Avoid it for ordinary CRUD where nobody will ever ask how a value changed.
CQRS: different models for writing and reading
Command Query Responsibility Segregation means the model that validates and accepts writes is not the model that serves reads. Writes go to a store designed for correctness (normalised, transactional); reads come from stores designed for each query (denormalised tables, a search index, a cache, a materialised feed).
You already use CQRS whenever you:
- keep a search index fed from the primary database,
- precompute a feed or a leaderboard,
- maintain a "summary" table updated asynchronously.
Cost: read models are eventually consistent with the write model, usually by milliseconds to seconds. The user who just made a change may not see it in the read model yet, so give them read-your-writes (read the write model for their own recent changes, or show the change optimistically). See data modelling for reads.
CQRS does not require event sourcing. Most real systems are CQRS on top of an ordinary database plus CDC.
Keeping read models in sync
The hard part of CQRS is updating the read models reliably. The naive approach, dual writes (the service writes the database, then updates the index or publishes an event), loses updates whenever the process crashes between the two, and can reorder them under concurrency.
Two reliable patterns:
Transactional outbox. In the same database transaction as the business change, insert a row into an outbox table. A relay reads the outbox in order and publishes to Kafka, marking rows sent. The event exists if and only if the change committed. Consumers must handle duplicates (the relay may publish twice), so make them idempotent.
Change data capture. Read the database’s own replication log (Postgres WAL, MySQL binlog) with a tool such as Debezium and publish each committed row change to Kafka, keyed by primary key so changes to one row stay in order. No application changes, nothing missed, and the stream reflects exactly what committed. The trade-off: events are row-level (what changed in the table), not business-level (why), unless you capture the outbox table itself, which combines both.
| Outbox | CDC on tables | |
|---|---|---|
| Event shape | business events you design | row changes |
| Application change | write the outbox row | none |
| Completeness | everything written through the app | everything that committed, including manual fixes |
| Typical use | publishing domain events to other services | feeding search, caches, warehouses |
Projections and rebuilding
A projection consumes the stream and writes a read model. Make it:
- Idempotent: process by key and version so replays and duplicates converge.
- Rebuildable: to change a read model, build a new one from the beginning of the log (or a snapshot plus the log) beside the old, then switch reads over.
- Monitored for lag: the time from commit to visible in the read model is your freshness promise.
This is exactly how Design Yelp keeps its search index in step with the business database.
In the interview
- Default: an ordinary transactional store as the source of truth, read models fed by CDC or an outbox, and read-your-writes where users notice.
- Reach for event sourcing when history matters: ledgers, order books, collaborative edits.
- Never propose dual writes without saying why they are safe (they rarely are).
Checklist
- What is the source of truth: current-state tables or an event log.
- Which read models exist, and what feeds them (CDC or outbox).
- How duplicates and reordering are handled in projections.
- The freshness lag of each read model, and read-your-writes for the writer.
- How a read model is rebuilt, and (if event-sourced) snapshots and event versioning.