SysDesignPrep.com
Study guide 88 of 183

"Kafka vs RabbitMQ vs SQS"

Choosing a messaging system: Kafka, RabbitMQ and Amazon SQS compared on model, ordering, replay, throughput, routing, delivery guarantees, delays and dead letters, operations and cost, with guidance on when to use a log, a broker or a managed queue.

Reading is half of it. See this used in a real interview: walk through Design a Distributed Message Queue →

"Why Kafka and not RabbitMQ?" is a frequent interview follow-up. The honest answer is that they solve different problems: Kafka is a distributed log for event streams, RabbitMQ is a flexible message broker for routing and task queues, and SQS is a fully managed queue that trades features for zero operations. Choosing well, and saying why in one sentence, signals real experience.

The models

Kafka: topics split into partitions, each an append-only log. Consumers track offsets and read at their own pace; messages stay for the retention period whether or not they were consumed. Multiple consumer groups each get the full stream. See how Kafka works.

RabbitMQ: producers publish to exchanges, which route messages to queues by rules (direct, topic patterns, fanout, headers). Consumers receive messages pushed from queues and acknowledge them; acknowledged messages are deleted. Rich per-message features. Newer RabbitMQ versions also offer replicated quorum queues and an append-only "streams" type.

SQS: a managed queue. Producers send, consumers poll, and a received message becomes invisible for a visibility timeout; if not deleted in time, it reappears for another consumer. Standard queues are nearly unlimited in throughput with best-effort ordering; FIFO queues give ordering per message group with lower throughput.

Comparison

KafkaRabbitMQSQS
Modelpartitioned logbroker with exchanges and queuesmanaged queue
Consumptionpull, offsetspush with acknowledgementspoll, visibility timeout
After consumptionretained until retention endsdeleted on ackdeleted by consumer
Replayyesno (except streams)no
Multiple independent consumersconsumer groups, each reads allfanout exchange to several queuesneeds SNS fanout to several queues
Orderingper partitionper queue (with one consumer)FIFO queues per message group
Throughputvery high (millions per second per cluster)high, lower than Kafkavery high (standard), limited per group (FIFO)
Routingby key to partitionsflexible patterns and headersnone (use SNS filtering)
Per-message delays and prioritiesnopriorities, delays with plugin or TTL tricksdelay up to 15 minutes
Dead lettersmanual patternbuilt inbuilt in (redrive policy)
Operationssubstantial (or use a managed service)moderatenone

When to use which

Kafka when:

  • Many consumers need the same events (analytics, search indexing, notifications, ML features).
  • You need replay: reprocess history after a bug fix, or bootstrap a new service.
  • Throughput is very high (clickstreams, logs, metrics, CDC).
  • Order per key matters at scale.

See Design an Ad Click Aggregator and change data capture.

RabbitMQ when:

  • You need task queues with per-message acknowledgement, retries and priorities.
  • Routing logic is complex (topic patterns, header-based routing).
  • Latency for individual messages matters more than bulk throughput.

SQS when:

  • You want a reliable queue with no operations, on AWS.
  • Workloads are task queues for background jobs and decoupling services, with variable volume. See background jobs.
  • Combined with SNS for fanout to several queues.

Delivery and duplicates

All three are at least once in normal use: consumers must be idempotent. Kafka offers exactly-once processing within Kafka through idempotent producers and transactions; SQS FIFO deduplicates sends within a five-minute window. Neither removes the need for idempotent side effects. See delivery semantics.

Work queue semantics on Kafka

Kafka can act as a job queue, but parallelism is capped by partition count, one slow message blocks its partition, and per-message retries and delays need extra topics. For job processing with uneven task durations, a traditional queue is usually simpler. See Design a Job Scheduler.

Other options

  • Google Pub/Sub and Azure Service Bus: managed equivalents with their own feature mixes.
  • Redis Streams: lightweight streams with consumer groups for smaller workloads. See Redis data structures.
  • Pulsar: log-based like Kafka with tiered storage and queue semantics built in.
  • NATS: very low-latency messaging, with JetStream for persistence.

In the interview

"Events go to Kafka, because several services consume them independently and we want replay; background tasks like sending emails use a queue (SQS) with visibility timeouts, retries and a dead-letter queue." That combination fits most designs, including Design a Notification System and Design a Message Queue.

Checklist

  • Log (Kafka) for event streams, many consumers, replay and high throughput.
  • Broker (RabbitMQ) for routing, priorities and per-message task handling.
  • Managed queue (SQS) for simple, operation-free task queues.
  • Idempotent consumers regardless of choice.
  • Dead-letter queues and retry policies for failed messages.

Open in your browser to sign in

Google does not allow sign-in inside this app's built-in browser. Open this page in Safari and sign in there. The link opens this same page.

Tap the ⋯ or share button at the top or bottom of the screen, then Open in browser. Or copy the link and paste it into Safari.