On this page
concept

Kafka

Created 2026-08-17 24 connections

Apache Kafka

Apache Kafka is a distributed, log-based event streaming platform designed for high-throughput, fault-tolerant, persistent data streams. In ecommerce, it serves as the event backbone connecting order management, inventory, personalisation, fraud detection, and analytics systems — replacing point-to-point API calls and overnight batch pipelines with real-time, decoupled event flows.


Core architecture

Kafka is a distributed log-based streaming platform that persists data on disk and leverages sequential I/O, enabling it to handle millions of messages per second with low latency. (Berd-i & Sons, 2026-03-20; The Code Wizard, unknown)

The core primitives:

  • Topic — a named, ordered, immutable log of events. Topics are organised logically based on business domain to simplify data management and retrieval. (Berd-i & Sons, 2026-03-20)
  • Partition — each topic is divided into partitions, enabling parallel consumption and horizontal scaling. Partition count is set at creation time and increasing it requires careful rebalancing.
  • Broker — a Kafka server node. A cluster of brokers stores all partitions and replicates them for fault tolerance.
  • Consumer Group — a set of consumers that jointly read a topic, each consuming from a distinct subset of partitions. This enables parallelism without duplicate processing.
  • Retention — Kafka retains messages on disk for a configurable period (or size), meaning consumers can replay history — a capability that distinguishes Kafka from traditional message queues which delete on acknowledgement.

By 2026, Event-Driven Architecture (EDA) has emerged as a cornerstone for modern software solutions, with businesses in ecommerce requiring systems that handle high volumes of data with low latency. (Berd-i & Sons, 2026-03-20)


Ecommerce use cases

Confluent's retail solutions page lists the following primary Kafka use cases in retail (Confluent, live as-of 2026):

  1. Inventory visibility across all sales channels — real-time inventory updates prevent stockouts and overselling.
  2. Customer 360 and personalised offers — streaming customer interaction data enables personalised offers and loyalty programmes.
  3. Abandoned cart triggers — capturing real-time user behaviour enables triggered re-engagement strategies.
  4. Fraud detection — streaming order and payment events through detection models before fulfilment.
  5. Demand forecasting — streaming pipelines feed real-time signals into forecasting models.
  6. Real-time replenishment — Kafka drives store replenishment decisions based on live sales signals.

Kai Waehner (2024-08-30) further identifies: supply chain and last-mile tracking, and edge-to-cloud streaming from physical store devices as core retail use cases.

Order processing decoupling

DZone's ecommerce migration article states that companies can replace monolithic order flows with independent services that exchange OrderPlaced, InventoryUpdated, and other events via Kafka topics. Kafka eliminates tight coupling between services — rather than calling each other directly via REST, services communicate by producing and consuming Kafka events. (DZone, unknown)

CDC and event streaming

Change Data Capture (CDC) with Kafka and Debezium turns databases into real-time event streams without modifying application code or adding polling overhead. By 2026, CDC is described as "the invisible plumbing beneath nearly every modern event-driven data pipeline." (The Backend Developers, 2026)

Debezium 2.5+ introduced improvements including better exactly-once semantics support, incremental snapshots without table locks, and native support for Kafka 4.0's KRaft mode (as-of 2026). (The Backend Developers, 2026; volatile)

Kai Waehner's 2024 analysis argues that by leveraging Kafka and Apache Flink together, unified commerce platforms can achieve high responsiveness, scalability, and seamless customer experience across online and offline channels. Jay Kreps (Confluent CEO) announced Confluent Cloud for Apache Flink reaching GA at Kafka Summit London 2024, positioning the Kafka + Flink + Apache Iceberg combination as a "Universal Data Product" — unifying operational and analytical data on a single streaming platform. (Kafka Summit London 2024 keynote, 2024-03-19)


Kafka vs alternatives

ToolTypeThroughputPersistenceReplayBest fit
Apache KafkaDistributed logMillions/secDisk✅ YesHigh-volume event streaming, event sourcing, CDC
RabbitMQMessage broker (AMQP)Tens of thousands/secMemory (optional disk)❌ NoComplex routing, task queues, low-latency delivery
Amazon KinesisManaged streaming (AWS)Scales managedDisk✅ YesAWS-native streaming, analytics
Google Pub/SubManaged streaming (GCP)Scales managedManagedLimitedGCP-native event distribution

(The Code Wizard, unknown; AWS in Plain English, unknown; Estuary, 2026)

RabbitMQ stores messages in memory (with optional persistence), while Kafka persists data to disk and enables message replay — a distinction that matters for ecommerce event sourcing scenarios where audit trails and re-processing are required. (The Code Wizard, unknown)

Estuary's 2026 comparison states that Kafka clearly dominates in high-throughput scenarios, especially when dealing with large-scale event streams such as telemetry or clickstream data. (Estuary, 2026)


Managed Kafka: 2026 landscape

Apache Kafka itself carries no software licence fee (open-source), but infrastructure, egress, and engineering hours push real production costs into six figures. Schema management, connectors, monitoring, and replication are the factors that often determine whether the free Kafka licence stays cheap or turns into a larger platform bill. (Airbyte, 2026)

Confluent Cloud (as-of 2026)

Charges per GB in/out plus compute. Elastic CKU model auto-scales compute based on demand — scaling is managed and abstracted from the operator, in contrast to self-managed Kafka. Private Network Interface (PNI) on AWS can cut cross-AZ costs by up to 50%. (AutoMQ, 2026; volatile)

Amazon MSK (as-of 2026)

MSK Express brokers offer approximately 20x faster scaling than standard MSK brokers (approximately 5 minutes vs 20–40 minutes), with unlimited pay-as-you-go storage. (AutoMQ, 2026; volatile)

Redpanda (as-of 2026)

Built in C++ (removing Java overhead and replication inefficiencies). Available in three deployment models: single-tenant Dedicated clusters, Bring Your Own Cloud (BYOC), and multi-tenant Serverless. (AutoMQ, 2026; volatile)

Redpanda cost advantage: AutoMQ's 2026 analysis cites vendor benchmarks claiming up to 6x lower cost than traditional Kafka deployments. (AutoMQ, 2026) — However, this benchmark originates from Redpanda's own vendor materials and is surfaced by AutoMQ, also a competing vendor. No independent third-party validation found. Should be treated as a marketing claim rather than a verified benchmark.

For both self-managed Kafka and Redpanda, scaling is a manual operation involving partition rebalancing and data copying, which for large clusters can take hours and materially impact performance — Confluent Cloud auto-scales. (AutoMQ, 2026)


At-scale implementations

Shopify

Shopify has run self-managed Kafka on Google Kubernetes Engine with brokers as Kubernetes StatefulSets with dedicated Persistent Volumes since 2014. (Factor House, unknown; Shopify Engineering blog)

Key scale figures (as-of 2024–2026, volatile):

  • 66 million Kafka messages per second peak during Black Friday / Cyber Monday. (Factor House, unknown; Medium/@shaikhjavedmail, unknown)
  • 1.75 trillion messages per month across self-managed clusters. (Shopify Engineering blog, referenced via Factor House)
  • 100,000 Debezium CDC records per second peak (65,000/sec sustained) from 100+ MySQL database shards. (Factor House, unknown)

Shopify's services produce events through Monorail, an internal event abstraction layer that enforces schemas, supports schema versioning, and defines explicit producer-consumer contracts on top of raw Kafka topics. (Factor House, unknown)

Shopify engineers contributed a lock-free snapshot mode for Debezium's MySQL connector to the open-source project, illustrating how at-scale retailers directly shape Kafka ecosystem tooling. (Factor House, unknown)

Kafka is highly sensitive to concurrent broker restarts, which can trigger terabytes of partition rebalancing; Shopify mitigates this with pod anti-affinity and readiness probe combinations. Shopify reduced cluster migration time from six months to two weeks by reusing persistent disks during Kubernetes upgrades. (Factor House, unknown; 2024-09-17 YouTube talk — low confidence, Apify unavailable)

Walmart (as-of 2024, stale-risk)

Walmart Kafka volume figures are inconsistent across sources. KSolves article (unknown date) states "approximately 8,500 Kafka nodes processing approximately 11 billion events per day." The Walmart Global Tech Blog on Medium (June 2024) states "trillions of messages per day" from a fleet serving 25,000+ consumers. The Confluent replenishment case study references "tens of billions of messages from 100 million SKUs." These likely reflect different time periods and different sub-systems (replenishment vs total platform vs CDP) — none should be quoted as a single headline figure without qualification. (KSolves, unknown; Walmart Global Tech Blog, 2024-06; Confluent, unknown)

Walmart's Kafka infrastructure maintained 99.99% availability while serving 25,000+ consumers across public and private clouds (as-of June 2024). (Walmart Global Tech Blog, 2024-06; volatile)

Walmart's real-time replenishment system processes more than tens of billions of messages from close to 100 million SKUs in less than three hours, at throughputs of 85 GB messages per minute, across 5,000+ stores. (Confluent, unknown; volatile)

Confluent's Walmart case study states that Walmart's Customer Data Platform ingests around 40 billion events per day. (Confluent, unknown; volatile)

Zalando (historical)

Zalando built Nakadi — an open-source REST abstraction layer over Kafka — as their central Event Bus Service, handling hundreds of gigabytes per second across 1,000+ event types, enabling their transition from a monolith to microservices. (Kai Waehner 2024-08-30 blog, referencing Zalando engineering blog)


Limitations and operational complexity

  • Engineering cost: Operational pain points for self-managed Kafka include cluster management, broker sizing, partition tuning, consumer lag monitoring, and the specialist engineering expertise required to keep it running reliably at production scale. (Airbyte, 2026)
  • Scaling is manual: Scaling self-managed Kafka requires partition rebalancing and data copying, which for large clusters can take hours and impact performance. (AutoMQ, 2026)
  • Broker restart sensitivity: Concurrent broker restarts can trigger terabytes of partition rebalancing. (Factor House, unknown)
  • Total cost of ownership: Despite open-source licensing, real production costs reach six figures once infrastructure, egress, and engineering hours are counted. Schema management, connectors, monitoring, and replication drive the bill upward. (Airbyte, 2026)

Key terms

TermMeaning
TopicNamed, ordered, immutable log of events on a specific subject (e.g. orders.placed)
PartitionSubdivision of a topic enabling parallel consumption and horizontal scaling
BrokerA Kafka server node that stores and serves partitions
Consumer GroupA set of consumers jointly reading a topic, each from a distinct partition subset
OffsetThe position of a consumer in a partition — how far it has read
RetentionHow long Kafka keeps messages on disk (enables replay; distinguishes Kafka from queues)
KRaft modeKafka's newer Raft-based consensus removing the ZooKeeper dependency (Kafka 4.0+)
MonorailShopify's internal event abstraction layer enforcing schemas on top of raw Kafka topics
NakadiZalando's open-source REST abstraction layer over Kafka as a central Event Bus Service
Managed KafkaKafka-as-a-service: Confluent Cloud, Amazon MSK, Redpanda Cloud

What practitioners report

No Reddit MCP available in this harvest (known recurring Cowork cloud gap). Practitioner signals sourced from trade press and engineering blogs only.

  • Shopify engineering (Factor House / Shopify Engineering blog) describes Kafka as "the data highway" between Shopify's sharded monolith, real-time processing layer, and data warehouse — running since 2014. Operational complexity management (Kubernetes anti-affinity, readiness probes, PVC reuse during upgrades) is a recurring theme in their engineering posts.
  • DZone's practitioner framing: companies migrating from monoliths replace synchronous REST chains with Kafka event flows, starting with OrderPlaced / InventoryUpdated events as the initial domain split.
  • The Backend Developers (2026): CDC is "the invisible plumbing beneath nearly every modern event-driven data pipeline" — positioning Kafka not as optional infrastructure but as the assumed backbone.

Dangling references (future frontiers)

  • Event-Driven Architecture — parent architectural pattern; no dedicated concept page
  • Outbox Pattern — atomically safe write + event emission pattern; closely related to Change Data Capture; no concept page
  • Apache Flink — stream processing complement to Kafka; Confluent Cloud GA 2024; no entity page
  • Kafka Streams — Kafka's built-in stream processing library (used in the ecommerce monolith-to-microservices migration case study); no concept page
  • Schema Registry — Confluent's schema enforcement layer referenced in Shopify Monorail context; no concept page
  • KRaft Mode — Kafka 4.0's ZooKeeper-free consensus mode; briefly referenced; no concept page
Research agent · 2026-08-17