SysDesignPrep.com
Study guide 89 of 183

"Synchronous vs asynchronous communication"

When services should call each other directly and when they should communicate through queues and events: latency, coupling, availability, consistency and debugging trade-offs, request-reply over queues, the 202 Accepted pattern, and how to split a user flow into sync and async parts.

Reading is half of it. See this used in a real interview: walk through Design DoorDash →

Every arrow in a design diagram is either a synchronous call (the caller waits for a response) or an asynchronous message (the caller hands off work and moves on). The choice affects latency, availability, consistency and how hard the system is to debug. Interviewers notice when every arrow is a synchronous call, because one slow service then slows everything. They also notice when everything goes through Kafka for no reason. This guide gives you the rule of thumb.

Synchronous calls

HTTP or gRPC request and response: the caller blocks until the callee answers (or times out).

Good:

  • Simple to write, reason about and debug: one request, one stack of calls, one trace.
  • The caller gets an immediate answer, including errors.
  • Natural for reads and for anything the user is waiting to see.

Costs:

  • Temporal coupling: the callee must be up and fast right now. Availability multiplies: a request that depends on five services at 99.9 % each is at best about 99.5 %.
  • Latency adds up along call chains, and the slowest dependency sets the pace. See tail latency.
  • Failures cascade: a slow dependency ties up the caller's threads. Timeouts, circuit breakers and bulkheads are mandatory. See rate limiting and resilience.

Asynchronous messaging

The caller publishes a message or event to a queue or log and continues; consumers process it later.

Good:

  • Decoupling in time: the consumer can be down or slow; messages wait.
  • Load levelling: bursts are absorbed by the queue and processed at a steady rate. See load shedding and backpressure.
  • Fan-out: many consumers react to one event without the producer knowing about them.
  • Retries happen in the background without the user waiting.

Costs:

  • Eventual consistency: the effect is not visible immediately.
  • Harder debugging: flows are spread across services and time; tracing must follow messages. See distributed tracing.
  • Duplicates and ordering need handling. See delivery semantics.
  • More infrastructure to run.

A rule of thumb

  • Synchronous for what the user needs right now to continue: reads, validation, and the minimal write that must succeed before you can say "done" (create the order, authorise the payment).
  • Asynchronous for everything that can happen after the response: notifications, emails, search indexing, analytics, thumbnails, recommendations updates, fraud review, payouts.

The user's request path becomes short and dependable; side effects happen reliably in the background.

Splitting a flow

Placing a food order:

StepModeWhy
Validate cart, price, addresssyncuser needs errors now
Authorise paymentsyncmust succeed before confirming
Create order record (plus outbox event)syncthe confirmation depends on it
Notify restaurantasynccan retry; the restaurant tablet may be offline
Dispatch courierasynctakes time; has its own retries
Send confirmation email and pushasyncnot needed for the response
Update analytics, recommendationsasynceventual is fine

The order is written together with an outbox event in one transaction, so the asynchronous steps are guaranteed to happen. See change data capture and the outbox pattern and Design DoorDash.

Long-running requests: 202 Accepted

When work takes seconds or minutes (video processing, report generation, large imports), do not hold the connection:

  1. Accept the request, enqueue the job, return 202 Accepted with a job id or status URL.
  2. The client polls the status, or receives a push, webhook or WebSocket message when done.

See background jobs and webhooks.

Request-reply over messaging

Sometimes you want asynchronous transport but still need a reply (a slow downstream, or traffic that must survive restarts). Send a request message with a correlation id and a reply topic; the responder publishes the answer there. Useful, but often a sign that a synchronous call with retries, or a workflow, would be simpler. See workflow orchestration.

Choreography and chains

Long chains of asynchronous events between many services (A emits, B reacts and emits, C reacts...) are flexible but hard to follow. When a process has many steps, branches or compensations, make it an explicit orchestrated workflow instead. See microservices.

In the interview

Mark arrows as sync or async in your diagram, and say: "The request path does only the work the user waits for; everything else is published as events and processed asynchronously with retries." In Design Instagram, uploading returns quickly while fan-out, thumbnails and notifications run in the background. In Design a Payment System, authorisation is synchronous; capture, ledger updates and payouts are asynchronous and idempotent.

Checklist

  • Sync for reads and the minimal write the user waits for.
  • Async for side effects, fan-out, slow work and anything retryable.
  • Timeouts, circuit breakers and bulkheads on every sync call.
  • Outbox for reliable handoff from sync to async.
  • 202 Accepted with status polling or push for long jobs.
  • Orchestration instead of long event chains.

Open in your browser to sign in

Google does not allow sign-in inside this app's built-in browser. Open this page in Safari and sign in there. The link opens this same page.

Tap the ⋯ or share button at the top or bottom of the screen, then Open in browser. Or copy the link and paste it into Safari.