"Kafka vs RabbitMQ vs SQS"
Choosing a messaging system: Kafka, RabbitMQ and Amazon SQS compared on model, ordering, replay, throughput, routing, delivery guarantees, delays and dead letters, operations and cost, with guidance on when to use a log, a broker or a managed queue.
Reading is half of it. See this used in a real interview: walk through Design a Distributed Message Queue →
"Why Kafka and not RabbitMQ?" is a frequent interview follow-up. The honest answer is that they solve different problems: Kafka is a distributed log for event streams, RabbitMQ is a flexible message broker for routing and task queues, and SQS is a fully managed queue that trades features for zero operations. Choosing well, and saying why in one sentence, signals real experience.
The models
Kafka: topics split into partitions, each an append-only log. Consumers track offsets and read at their own pace; messages stay for the retention period whether or not they were consumed. Multiple consumer groups each get the full stream. See how Kafka works.
RabbitMQ: producers publish to exchanges, which route messages to queues by rules (direct, topic patterns, fanout, headers). Consumers receive messages pushed from queues and acknowledge them; acknowledged messages are deleted. Rich per-message features. Newer RabbitMQ versions also offer replicated quorum queues and an append-only "streams" type.
SQS: a managed queue. Producers send, consumers poll, and a received message becomes invisible for a visibility timeout; if not deleted in time, it reappears for another consumer. Standard queues are nearly unlimited in throughput with best-effort ordering; FIFO queues give ordering per message group with lower throughput.
Comparison
| Kafka | RabbitMQ | SQS | |
|---|---|---|---|
| Model | partitioned log | broker with exchanges and queues | managed queue |
| Consumption | pull, offsets | push with acknowledgements | poll, visibility timeout |
| After consumption | retained until retention ends | deleted on ack | deleted by consumer |
| Replay | yes | no (except streams) | no |
| Multiple independent consumers | consumer groups, each reads all | fanout exchange to several queues | needs SNS fanout to several queues |
| Ordering | per partition | per queue (with one consumer) | FIFO queues per message group |
| Throughput | very high (millions per second per cluster) | high, lower than Kafka | very high (standard), limited per group (FIFO) |
| Routing | by key to partitions | flexible patterns and headers | none (use SNS filtering) |
| Per-message delays and priorities | no | priorities, delays with plugin or TTL tricks | delay up to 15 minutes |
| Dead letters | manual pattern | built in | built in (redrive policy) |
| Operations | substantial (or use a managed service) | moderate | none |
When to use which
Kafka when:
- Many consumers need the same events (analytics, search indexing, notifications, ML features).
- You need replay: reprocess history after a bug fix, or bootstrap a new service.
- Throughput is very high (clickstreams, logs, metrics, CDC).
- Order per key matters at scale.
See Design an Ad Click Aggregator and change data capture.
RabbitMQ when:
- You need task queues with per-message acknowledgement, retries and priorities.
- Routing logic is complex (topic patterns, header-based routing).
- Latency for individual messages matters more than bulk throughput.
SQS when:
- You want a reliable queue with no operations, on AWS.
- Workloads are task queues for background jobs and decoupling services, with variable volume. See background jobs.
- Combined with SNS for fanout to several queues.
Delivery and duplicates
All three are at least once in normal use: consumers must be idempotent. Kafka offers exactly-once processing within Kafka through idempotent producers and transactions; SQS FIFO deduplicates sends within a five-minute window. Neither removes the need for idempotent side effects. See delivery semantics.
Work queue semantics on Kafka
Kafka can act as a job queue, but parallelism is capped by partition count, one slow message blocks its partition, and per-message retries and delays need extra topics. For job processing with uneven task durations, a traditional queue is usually simpler. See Design a Job Scheduler.
Other options
- Google Pub/Sub and Azure Service Bus: managed equivalents with their own feature mixes.
- Redis Streams: lightweight streams with consumer groups for smaller workloads. See Redis data structures.
- Pulsar: log-based like Kafka with tiered storage and queue semantics built in.
- NATS: very low-latency messaging, with JetStream for persistence.
In the interview
"Events go to Kafka, because several services consume them independently and we want replay; background tasks like sending emails use a queue (SQS) with visibility timeouts, retries and a dead-letter queue." That combination fits most designs, including Design a Notification System and Design a Message Queue.
Checklist
- Log (Kafka) for event streams, many consumers, replay and high throughput.
- Broker (RabbitMQ) for routing, priorities and per-message task handling.
- Managed queue (SQS) for simple, operation-free task queues.
- Idempotent consumers regardless of choice.
- Dead-letter queues and retry policies for failed messages.