- Published on
Apache Kafka or RabbitMQ for My Case?
- Authors

- Name
- Mehdi Akiki
Most articles comparing Kafka and RabbitMQ are too generic to be useful. They usually say "it depends," list a few features, and stop there. That is not enough.
The real question is this:
What is the thing your system is trying to preserve?
If the answer is "a durable history of events that many systems may need to read now or later", Kafka is usually the better fit. Kafka is built as a distributed event streaming platform, and consumers can control and even rewind their offsets to re-consume data. Kafka also retains events according to retention settings, potentially for a very long time.
If the answer is "a task that should be handed to a worker, acknowledged, retried if needed, and then forgotten", RabbitMQ is usually the better fit. RabbitMQ's queues are built around delivery, acknowledgements, routing, and worker-style messaging. Durable replicated quorum queues are now RabbitMQ's recommended HA queue type, and RabbitMQ also supports streams when you want more log-like behavior.
That distinction is the whole game.
What Kafka is really for
Kafka is strongest when your system revolves around events as a source of truth.
A producer writes events to a topic. Kafka stores them durably, partitions them for scale, and lets many consumers read them independently. Consumers track offsets, not the broker deleting messages after a single successful delivery. That means one consumer can build a search index, another can power alerts, another can generate analytics, and a fourth can reprocess historical data after a bug fix. Kafka's design is built around retained logs, replay, partitioned scaling, and stream processing. Modern Kafka also runs in KRaft mode, so current deployments no longer need ZooKeeper.
That is why Kafka is a natural fit for systems like:
- telemetry pipelines
- observability backbones
- audit/event histories
- clickstreams
- security event ingestion
- event sourcing
- multi-consumer data platforms
What RabbitMQ is really for
RabbitMQ is strongest when your system revolves around message delivery and work dispatch.
RabbitMQ gives you exchanges, bindings, queues, acknowledgements, and flexible routing patterns. It supports several protocols, including AMQP and MQTT, and its queues are FIFO by default. If you want to hand a message to a worker, wait for acknowledgement, retry on failure, route to a dead-letter flow, or implement classic background-job semantics, RabbitMQ is excellent.
That makes RabbitMQ a great fit for things like:
- email sending jobs
- PDF generation
- image processing
- webhook dispatch
- workflow steps
- request/response integration
- service-to-service task handoff
- complex routing topologies
RabbitMQ has added streams, which narrows the gap in some cases, but its center of gravity is still different. Even RabbitMQ's own docs frame streams as a separate persistent replicated structure with different storage and consumption semantics from queues.
For my case: Kafka fits better
For this kind of product, RabbitMQ is usually the wrong primary backbone.
If you are ingesting SDK or agent data from customer systems, then the raw incoming event stream is not just a temporary message. It is the product's raw material.
You usually want all of this:
- absorb bursts without dropping the system
- retain events for investigation
- replay them when a parser or enrichment step is fixed
- feed multiple downstream consumers independently
- build analytics later from data collected earlier
- support debugging and "what happened?" workflows
- isolate slow consumers from fast producers
- preserve ordered processing at least within a key such as tenant, run, session, or trace
That is Kafka territory. Kafka consumers can re-read from offsets, topics retain data according to policy, and partitions give you scalable parallelism with ordering guarantees per partition.
For an observability or agent-tracing SaaS, the event stream is not merely transport. It is the system record.
RabbitMQ can move those messages, yes. But if later you need replay, multi-consumer fan-out at scale, retrospective analysis, backfills, re-indexing, billing recomputation, or new downstream products built on old data, RabbitMQ starts feeling like the wrong primitive. The discussion circles around exactly that difference: queue semantics for work dispatch versus log semantics for retained event history.
A practical way to think about it
Ask one brutal question:
If a consumer is broken for six hours, do I want the data to still be there so I can catch up and maybe replay it next week?
If the answer is yes, lean Kafka. Kafka is explicitly designed to store events durably and let consumers resume from their own positions.
Ask a second question:
Do I mostly want to hand one piece of work to one worker and know whether it succeeded or needs retrying?
If the answer is yes, lean RabbitMQ. RabbitMQ's ack model, queue semantics, and routing are excellent for that.
Your case sounds much more like the first than the second.
What I would choose for your architecture
I would use Kafka as the ingestion spine.
A clean design would look like this:
SDK/agent -> ingestion API -> Kafka -> multiple consumers
Then those consumers can independently do things like:
- normalize raw events
- enrich traces with metadata
- power live debugging views
- write searchable storage
- compute usage/billing
- trigger alert rules
- archive raw history
- run offline analytics or backfills
That architecture matches Kafka's natural strengths: retained log, many independent consumers, replay, partitioned scale, and durable history.
Where RabbitMQ still helps
This does not mean RabbitMQ is useless in your stack.
It can still be the right tool for:
- internal background jobs
- one-off command execution
- email/report generation
- webhook retries
- admin tasks
- orchestrating finite workflow steps
In other words:
- Kafka for product data
- RabbitMQ for operational jobs
That split is often cleaner than trying to force one broker to do everything.
The mistake people make
The common mistake is picking RabbitMQ because the first version of the system only has a few services and low traffic.
That can be fine for a pure job system.
But for an event product, the hard part is not today's throughput. The hard part is what you will need in six months:
- replay
- backfill
- auditability
- multiple downstream products
- historical analysis
- debugging from raw truth
- safe decoupling between producers and consumers
If those are on the roadmap, starting with Kafka is usually the more honest design.
The other mistake is repeating outdated benchmark numbers like "RabbitMQ does X and Kafka does Y messages per second." That is sloppy engineering. Actual throughput depends on message size, durability settings, replication, batching, consumer behavior, queue/topic design, disk, network, and operational choices. Even the sources you find online are useful more for framing the tradeoff than for trusting old fixed performance numbers.
Final recommendation
For this kind of use case, I would choose Apache Kafka as the primary event backbone.
Use it if your product needs:
- durable ingestion
- replay
- multi-consumer fan-out
- historical debugging
- analytics on past events
- event-driven product features
- scalable tenant/session/trace partitioning
Use RabbitMQ only for the parts of the system that are truly about tasks, not event history.
So the real answer is:
If you are building an observability / agent-event / telemetry SaaS, choose Kafka first. If you are dispatching background work, choose RabbitMQ first. If you are doing both, use both, but do not put RabbitMQ in charge of your product's core event history.
I build and scale reliable production systems. Open to full-time and freelance work with U.S.-based teams that value ownership and execution.
Got something in mind?
Book a Discovery Call