Event-driven architecture
Designing systems around events: events versus commands, notification events versus event-carried state transfer, topics and schemas, choreography, eventual consistency, idempotent consumers, ordering, replay, event storming, and the pitfalls that make event-driven systems hard to operate.
Reading is half of it. See this used in a real interview: walk through Design DoorDash →
In an event-driven architecture, services announce facts ("order placed", "payment captured", "user signed up") and other services react, instead of calling each other directly. It decouples teams, absorbs load spikes and makes it easy to add new behaviour without touching the producer. It also makes flows harder to follow and introduces eventual consistency everywhere. Interviewers like candidates who use events where they help and can name the costs.
Events versus commands
- An event describes something that happened:
OrderPlaced. The producer does not know or care who listens. Past tense, immutable. - A command asks for something to happen:
ChargeCard. It has an intended recipient and can be rejected.
Mixing them up causes confusion: an "event" named SendEmail is really a command, coupling the producer to one consumer.
Kinds of events
| Kind | Payload | Trade-off |
|---|---|---|
| Notification | ids only (OrderPlaced {order_id}) | small and stable, but consumers must call back for details, adding coupling and load |
| Event-carried state transfer | the relevant data (OrderPlaced {order_id, items, total, customer}) | consumers act without calling back and can build their own copies; larger events, schema discipline needed |
| Domain events in event sourcing | every state change, as the source of truth | full history and replay; most complex. See event sourcing and CQRS |
Event-carried state transfer is usually the better default for cross-service integration.
Plumbing
- A log or broker: Kafka for durable, replayable, high-volume streams; a queue or managed pub/sub for simpler cases. See Kafka vs RabbitMQ vs SQS.
- Topics organised by domain (
orders,payments), partitioned by the entity id whose order matters. - Schemas in a registry with compatibility rules, because events are a public contract. See schema evolution and serialization.
- Reliable publishing with the outbox pattern, so the database change and the event cannot diverge. See change data capture and the outbox pattern.
- Envelope metadata: event id, type, version, timestamp, source, correlation and trace ids.
Consumers
- Idempotent: deliveries are at least once. See delivery semantics.
- Order-aware: ordering is per partition only; handle out-of-order events with versions or sequence numbers.
- Own their data: consumers build local read models from events instead of querying the producer.
- Dead-letter queues for poison events, with replay tooling.
Benefits
- Loose coupling: add a fraud check, a loyalty program or an analytics pipeline by subscribing, with no change to the order service.
- Resilience: a consumer outage delays its work but does not fail the producer.
- Load levelling: bursts are absorbed by the log. See load shedding and backpressure.
- Replay: rebuild a read model or bootstrap a new service from history.
- Audit trail of what happened.
Costs and pitfalls
- Eventual consistency: after placing an order, the order history page may not show it yet. Design the UI for it (show the user's own recent writes). See consistency models.
- Invisible flows: no single place shows the end-to-end process. Use tracing across events and, for multi-step processes, explicit orchestration. See distributed tracing and workflow orchestration.
- Event spaghetti: long chains where services react to each other's reactions, with cycles and surprises.
- Schema coupling: consumers depend on event fields; breaking changes ripple widely.
- Debugging and testing need tooling: event browsers, replays, contract tests. See testing distributed systems.
Designing events
Event storming is a workshop technique: list the domain events on a timeline ("cart checked out", "payment authorised", "order accepted by restaurant", "courier assigned"), then find the commands that cause them, the aggregates that own them, and the boundaries between services. It produces both the event catalogue and the service boundaries. See microservices.
Where it fits
- Side effects of a core transaction: notifications, search indexing, analytics, recommendations. See Design a Notification System.
- Integration between teams' services in a large organisation.
- High-volume pipelines: clicks, location updates, metrics.
- Less suited to: request-response interactions where the user needs an immediate answer, or tightly coupled multi-step transactions (use orchestration with explicit compensation instead).
In the interview
"The order service writes the order and an outbox event in one transaction; OrderPlaced carries the order details on a Kafka topic partitioned by order id; payments, restaurant notification, analytics and search consume it independently and idempotently; the multi-step fulfilment flow itself runs as an orchestrated workflow." See Design DoorDash and Design Instagram.
Checklist
- Events are past-tense facts; commands are requests.
- Event-carried state transfer for integration; schemas in a registry.
- Outbox for reliable publishing; partitioning by entity id.
- Idempotent, order-aware consumers with dead-letter queues.
- Tracing across events; orchestration for long processes.
- Eventual consistency designed into the user experience.