Workflow orchestration and durable execution
Running multi-step processes reliably: choreography versus orchestration, durable execution engines like Temporal, state machines, retries and timeouts per step, compensation, long waits and human steps, data pipeline schedulers like Airflow, and versioning running workflows.
Reading is half of it. See this used in a real interview: walk through Design DoorDash →
Many business processes span several services and minutes or days: place an order, charge the card, notify the restaurant, assign a courier, wait for delivery, pay the courier. Each step can fail, time out or need a retry, and the process must survive crashes and deploys halfway through. Hand-written chains of queues and status columns become hard to follow and easy to break. Workflow orchestration makes the process explicit and durable. It is a strong answer whenever an interview design has a multi-step flow with failures.
Choreography versus orchestration
- Choreography: services react to each other's events (OrderPlaced triggers payment; PaymentSucceeded triggers dispatch). Loosely coupled, but the overall process exists only implicitly, spread across services, and is hard to observe or change.
- Orchestration: a coordinator (the workflow) calls each step in order and decides what happens next. The process is in one place, easy to read, monitor and modify; the orchestrator becomes a central dependency.
Simple flows with few steps suit choreography. Long, branching flows with compensation suit orchestration. See event sourcing and CQRS and microservices.
Durable execution
Engines such as Temporal (and its predecessor Cadence), AWS Step Functions and others let you write the workflow as code or a state machine, and they make it durable:
- Every step's result is recorded in a history (event sourcing of the workflow).
- If the worker crashes, another worker replays the history to restore the workflow's state exactly and continues from where it stopped.
- Timers can wait for hours or months ("wait 7 days, then send a reminder") without holding resources.
- Signals deliver external events (a courier accepted, a human approved).
Steps that call other systems are activities: they are retried according to a policy and must be idempotent, because a crash after the call but before the result is recorded leads to a retry. See delivery semantics.
Designing a workflow
For each step, decide:
- Timeout: how long before giving up on an attempt.
- Retry policy: backoff, maximum attempts, which errors are retryable.
- Compensation: what to undo if a later step fails (refund the payment, release the hold). This is the saga pattern. See distributed transactions and idempotency.
- Idempotency key: usually the workflow id plus step name.
Example, food delivery: authorise payment, send order to restaurant (wait for acceptance with a timeout), dispatch courier (retry matching until assigned), wait for delivery (signal), capture payment, pay out. If the restaurant rejects, void the authorisation and notify the customer. See Design DoorDash.
Human steps and long waits
Workflows can wait for people: manual review of a flagged payment, a host accepting a booking request within 24 hours. The workflow sleeps durably until a signal arrives or a timer expires, then continues or compensates. See Design Airbnb and bookings and reservations.
Determinism and versioning
Replay requires workflow code to be deterministic: no direct random numbers, clock reads or network calls inside the workflow logic (those go through the engine or activities). Changing workflow code while instances are running needs versioning, so old instances replay with the logic they started with. Plan for this; workflows that run for weeks will span many deploys.
Data pipeline orchestration
Batch data pipelines use schedulers such as Airflow or Dagster: a DAG of tasks (extract, transform, load, validate) run on a schedule, with dependencies, retries, backfills for past dates and alerts. Tasks should be idempotent per date partition, so reruns overwrite rather than duplicate. See data lakes and lakehouses and batch and stream processing.
When a status column is enough
A two-step process with one retry does not need an engine: a state column, a job queue and a periodic sweeper for stuck items work fine. Reach for orchestration when steps multiply, compensation appears, waits get long, or nobody can explain the current flow from the code. See background jobs.
Observability
Orchestration engines show every running workflow, its current step, history and failures, which is a big operational advantage: support staff can see exactly where an order is stuck, and engineers can retry or terminate instances. Track per-step latency and failure rates and alert on workflows stuck past their expected duration.
In the interview
For Design a Payment System or any multi-service flow: "the process runs as a durable workflow; each step is an idempotent activity with its own timeout and retry policy; failures trigger compensations; long waits use durable timers; and we can see every order's current state." It answers reliability, consistency and operability at once.
Checklist
- Orchestration for long, branching processes; choreography for simple reactions.
- Durable execution with history replay, timers and signals.
- Per-step timeouts, retry policies and idempotency keys.
- Compensations for partial failure (sagas).
- Deterministic workflow code; versioning for running instances.
- DAG schedulers with idempotent, partitioned tasks for data pipelines.
- Visibility into every running instance.