"Synchronous vs asynchronous communication"
When services should call each other directly and when they should communicate through queues and events: latency, coupling, availability, consistency and debugging trade-offs, request-reply over queues, the 202 Accepted pattern, and how to split a user flow into sync and async parts.
Reading is half of it. See this used in a real interview: walk through Design DoorDash →
Every arrow in a design diagram is either a synchronous call (the caller waits for a response) or an asynchronous message (the caller hands off work and moves on). The choice affects latency, availability, consistency and how hard the system is to debug. Interviewers notice when every arrow is a synchronous call, because one slow service then slows everything. They also notice when everything goes through Kafka for no reason. This guide gives you the rule of thumb.
Synchronous calls
HTTP or gRPC request and response: the caller blocks until the callee answers (or times out).
Good:
- Simple to write, reason about and debug: one request, one stack of calls, one trace.
- The caller gets an immediate answer, including errors.
- Natural for reads and for anything the user is waiting to see.
Costs:
- Temporal coupling: the callee must be up and fast right now. Availability multiplies: a request that depends on five services at 99.9 % each is at best about 99.5 %.
- Latency adds up along call chains, and the slowest dependency sets the pace. See tail latency.
- Failures cascade: a slow dependency ties up the caller's threads. Timeouts, circuit breakers and bulkheads are mandatory. See rate limiting and resilience.
Asynchronous messaging
The caller publishes a message or event to a queue or log and continues; consumers process it later.
Good:
- Decoupling in time: the consumer can be down or slow; messages wait.
- Load levelling: bursts are absorbed by the queue and processed at a steady rate. See load shedding and backpressure.
- Fan-out: many consumers react to one event without the producer knowing about them.
- Retries happen in the background without the user waiting.
Costs:
- Eventual consistency: the effect is not visible immediately.
- Harder debugging: flows are spread across services and time; tracing must follow messages. See distributed tracing.
- Duplicates and ordering need handling. See delivery semantics.
- More infrastructure to run.
A rule of thumb
- Synchronous for what the user needs right now to continue: reads, validation, and the minimal write that must succeed before you can say "done" (create the order, authorise the payment).
- Asynchronous for everything that can happen after the response: notifications, emails, search indexing, analytics, thumbnails, recommendations updates, fraud review, payouts.
The user's request path becomes short and dependable; side effects happen reliably in the background.
Splitting a flow
Placing a food order:
| Step | Mode | Why |
|---|---|---|
| Validate cart, price, address | sync | user needs errors now |
| Authorise payment | sync | must succeed before confirming |
| Create order record (plus outbox event) | sync | the confirmation depends on it |
| Notify restaurant | async | can retry; the restaurant tablet may be offline |
| Dispatch courier | async | takes time; has its own retries |
| Send confirmation email and push | async | not needed for the response |
| Update analytics, recommendations | async | eventual is fine |
The order is written together with an outbox event in one transaction, so the asynchronous steps are guaranteed to happen. See change data capture and the outbox pattern and Design DoorDash.
Long-running requests: 202 Accepted
When work takes seconds or minutes (video processing, report generation, large imports), do not hold the connection:
- Accept the request, enqueue the job, return 202 Accepted with a job id or status URL.
- The client polls the status, or receives a push, webhook or WebSocket message when done.
See background jobs and webhooks.
Request-reply over messaging
Sometimes you want asynchronous transport but still need a reply (a slow downstream, or traffic that must survive restarts). Send a request message with a correlation id and a reply topic; the responder publishes the answer there. Useful, but often a sign that a synchronous call with retries, or a workflow, would be simpler. See workflow orchestration.
Choreography and chains
Long chains of asynchronous events between many services (A emits, B reacts and emits, C reacts...) are flexible but hard to follow. When a process has many steps, branches or compensations, make it an explicit orchestrated workflow instead. See microservices.
In the interview
Mark arrows as sync or async in your diagram, and say: "The request path does only the work the user waits for; everything else is published as events and processed asynchronously with retries." In Design Instagram, uploading returns quickly while fan-out, thumbnails and notifications run in the background. In Design a Payment System, authorisation is synchronous; capture, ledger updates and payouts are asynchronous and idempotent.
Checklist
- Sync for reads and the minimal write the user waits for.
- Async for side effects, fan-out, slow work and anything retryable.
- Timeouts, circuit breakers and bulkheads on every sync call.
- Outbox for reliable handoff from sync to async.
- 202 Accepted with status polling or push for long jobs.
- Orchestration instead of long event chains.