"At-most-once, at-least-once and exactly-once delivery"
What message delivery guarantees really mean: why exactly-once delivery is impossible over a network but exactly-once processing is achievable, acknowledgements and retries, idempotent consumers, deduplication, transactional processing in Kafka, and ordering.
Reading is half of it. See this used in a real interview: walk through Design a Distributed Message Queue →
"Does the queue guarantee exactly-once delivery?" is one of the most common follow-up questions in system design interviews, and a trap. Over an unreliable network, a sender can never be sure a message arrived, so it must either risk losing it or risk sending it twice. What systems actually offer is a choice of trade-offs, plus techniques that make duplicates harmless. Understanding this lets you answer precisely instead of hand-waving.
The three guarantees
| Guarantee | Behaviour | How | Use when |
|---|---|---|---|
| At most once | never duplicated, may be lost | send without retry; acknowledge before processing | metrics samples, telemetry where loss is acceptable |
| At least once | never lost, may be duplicated | retry until acknowledged; acknowledge after processing | the default for almost everything |
| Exactly once (effectively) | each message affects state once | at-least-once delivery plus idempotency or transactions | payments, billing, counts that matter |
Why duplicates are unavoidable
A consumer processes a message, then crashes before its acknowledgement reaches the broker. The broker cannot tell this apart from a consumer that crashed before processing, so it redelivers. Likewise, a producer that times out waiting for an acknowledgement does not know whether the broker stored the message, so it retries. Any system that never loses messages will sometimes deliver them twice.
The acknowledgement point
When the consumer acknowledges (or commits its offset) determines the guarantee:
- Commit before processing: a crash after the commit loses the message (at most once).
- Commit after processing: a crash before the commit repeats the message (at least once).
Almost always choose the second and make processing safe to repeat.
Making duplicates harmless
Exactly-once processing = at-least-once delivery + idempotent effects:
- Natural idempotency: "set status to shipped" can run twice safely; "add 10 to balance" cannot.
- Idempotency keys: each message carries a unique id; the consumer records processed ids (with a unique constraint) in the same transaction as its state change, and skips ids it has seen. See distributed transactions and idempotency.
- Versioned or conditional writes: apply only if the stored version is older than the message's version.
- Upserts keyed by a deterministic id: writing the same aggregate for (item, window) twice gives the same row.
The deduplication record must be atomic with the effect. Checking "seen?" in Redis and then writing to Postgres leaves a window where a crash causes a duplicate or a loss.
Exactly once in Kafka
Kafka provides:
- Idempotent producers: the broker deduplicates retried sends within a partition using producer ids and sequence numbers.
- Transactions: a consume-process-produce job can write its output messages and commit its input offsets atomically, so downstream readers (with read-committed isolation) see each result once.
This gives exactly-once semantics inside Kafka. As soon as the effect leaves Kafka (a database write, an email, a charge on a card), you are back to idempotency at that boundary. Stream processors such as Flink extend this with checkpoints and transactional sinks. See message queues and streams.
Side effects in the outside world
Some effects cannot be undone or deduplicated by you: an email sent, a push notification, a request to a payment provider. Strategies:
- Pass an idempotency key to the provider (payment APIs support this).
- Record intent before acting and outcome after, so a retry can check what happened. See payments and ledgers.
- Accept rare duplicates where harmless (a duplicate notification), with deduplication by key over a time window to make them rarer. See Design a Notification System.
Ordering
Ordering interacts with retries:
- Brokers usually guarantee order within a partition (Kafka) or message group (SQS FIFO), not globally. Partition by the key whose order matters (account id, order id).
- Retrying a failed message while later ones proceed breaks order. If order matters, block the partition (or that key) until the message succeeds, or send failures to a dead-letter queue and design consumers to tolerate gaps.
- Consumers can enforce order with sequence numbers per key, discarding or buffering out-of-order messages.
Dead letters and poison messages
A message that always fails would be retried forever and block its partition. After a few attempts with backoff, move it to a dead-letter queue, alert, and provide tooling to inspect and replay it after a fix. See background jobs.
In the interview
Say: "Delivery is at least once; the consumer records the message id in the same transaction as its write, so processing is effectively exactly once. Order is guaranteed per key by partitioning on account id. Failures retry with backoff and go to a dead-letter queue." That covers the follow-ups for Design a Message Queue, Design a Payment System and Design an Ad Click Aggregator.
Checklist
- At-least-once delivery by default; acknowledge after processing.
- Idempotent consumers with deduplication atomic with the effect.
- Idempotent producers and transactions inside Kafka; idempotency keys at external boundaries.
- Ordering per key via partitioning; a clear retry policy that respects it.
- Dead-letter queues with replay tooling.