On this page
- How it works
- Poison messages vs transient failures
- Retry patterns before the DLQ
- Blocking (synchronous) retry
- Non-blocking (asynchronous) retry with backoff
- DLQ in ecommerce: key failure scenarios
- Order event processing
- Payment processing
- Inventory and fulfilment events
- The governance problem
- Configuration parameters
- AWS SQS (as-of 2026-06-08)
- When NOT to use a DLQ (AWS docs, 2026-06-08)
- Replay (redrive) discipline
- DLQ and related patterns
- Key terms
- Benchmarks / operational data (as-of 2026-08-21)
- What practitioners report
Dead Letter Queue (DLQ)
Dead Letter Queue (DLQ)
A dead letter queue (DLQ) is a special message queue that stores messages a system cannot process after exhausting its retry budget. In ecommerce, DLQs are the last line of defence in event-driven order pipelines: when a payment webhook, inventory update, or fulfilment event fails processing, the DLQ catches it rather than letting it be silently dropped or cause a consumer to block indefinitely.
How it works
A DLQ operates alongside a regular message queue using a redrive policy — a set of rules (chiefly maxReceiveCount) that govern when a message is moved out of the main queue into the dead letter queue. The message's failure reason is typically one of (AWS "What is a DLQ?" documentation, updated 2026-06-08):
- Erroneous message content — hardware, software, or network corruption changes the payload; receiver rejects it
- Changes in receiver's system — consumer API schema changes or the entity the message references no longer exists
- TTL expiry — the message's time-to-live expires before the consumer can process it
- Message size limit exceeded — payload exceeds the queue's maximum message size
Messages remain in the DLQ until developers inspect, fix the root cause, and redrive them back to the source queue (AWS docs, 2026-06-08).
Poison messages vs transient failures
Two distinct failure classes determine the correct DLQ handling strategy (Redpanda / Dunith Danushka, 2022; widely cited in 2024–2026 practitioner writing):
| Failure class | Characteristics | DLQ handling |
|---|---|---|
| Non-transient (poison pill) | Deterministic — always fails regardless of how many retries. Caused by malformed JSON, schema mismatches, validation failures, consumer code bugs | Route directly to DLQ; do not retry — retrying wastes resources and never succeeds |
| Transient | Non-deterministic — may self-heal. Caused by temporary downstream unavailability, network timeouts, database locks | Retry with exponential backoff first; only park in DLQ if all retry attempts are exhausted |
The Redpanda source (Dunith Danushka) was published November 2022. The taxonomy of poison-pill vs transient failures remains standard in 2026 practitioner writing, but Spring Kafka API specifics (ErrorHandlingDeserializer versions) may have changed. Treat code examples as illustrative, not current.
Retry patterns before the DLQ
Practitioners distinguish between two retry approaches before a message reaches the DLQ (Redpanda blog, 2022; Codemia knowledge hub):
Blocking (synchronous) retry
Consumer thread is suspended while the same message is retried. Drawback: clogs the consumer thread; prevents offset commit; all downstream messages from the same partition are held behind the failing one (Redpanda, 2022).
Non-blocking (asynchronous) retry with backoff
Failed messages are forwarded to one or more retry topics (separate from the DLQ), processed by a dedicated retry consumer. This frees the main consumer thread to continue with fresh messages.
A common Spring Kafka pattern (Redpanda, 2022) uses @RetryableTopic with exponential backoff across multiple retry topics before the message finally lands in the DLQ:
orders → orders-retry-0 (1s delay) → orders-retry-1 (2s delay) → orders-retry-2 (4s delay) → orders-dlt
In Kafka-based systems, the DLQ is typically a dead letter topic (DLT) — a regular topic used as the DLQ terminal state (Redpanda, 2022).
DLQ in ecommerce: key failure scenarios
Order event processing
Amazon's order processing uses DLQs to quarantine corrupted events that would otherwise trigger cascading failures across inventory, payment, and shipping services (referenced in designgurus.substack.com, 2025–2026 writeup).
Payment processing
Without a DLQ, a single malformed payment event in a Kafka partition blocks all subsequent payments from that partition. A cited case: a consumer retried a bad message >10,000 times over six hours, blocking every payment from that partition — DLQ would have isolated the bad message immediately (designgurus.substack.com, 2025–2026).
Inventory and fulfilment events
A Black Friday incident described by syarif (Level Up Coding, August 2026): a downstream inventory service went down for 40 minutes; 6,200 order events failed with no retry path. Post-incident, a Kafka DLQ was added to capture failed events rather than dropping them.
The governance problem
The most detailed operational account in 2026 practitioner writing (syarif / Level Up Coding, August 2026) documents a DLQ that became a liability:
- Kafka DLQ (
orders.dlq) was added after Black Friday incident with a three-line PR description: "adds DLQ for resilience." - No runbook, no owner, no retention policy, no scheduled review was assigned at creation
- 14 months later: 2.3 million messages had accumulated, some going back 14 months
- A compliance audit found the DLQ contained customer email addresses, shipping details, and partial payment metadata — data categories the team's retention policy (71 days) applied to
orders.createdbut had never been applied toorders.dlq
This case surfaces a structural risk: a DLQ is a data store, not just an error buffer. It inherits the PII and retention obligations of the messages it receives, but is rarely treated as such at creation time.
Source: syarif, "Before You Add a Dead Letter Queue," Level Up Coding / Medium, 2026-08-19
Configuration parameters
AWS SQS (as-of 2026-06-08)
- maxReceiveCount — number of times a message is received before being sent to the DLQ; optimize to avoid sending transiently-failing messages to DLQ too early (AWS docs)
- Redrive policy — rules defining source queue ↔ DLQ interaction; includes maxReceiveCount
- Redrive-to-source — SQS feature to replay DLQ messages back to their original queue (or a different destination) on demand (AWS docs)
- For FIFO queues: the DLQ must also be a FIFO queue to preserve message ordering (AWS docs)
When NOT to use a DLQ (AWS docs, 2026-06-08)
- When messages must be retried indefinitely (e.g., waiting for a dependent service to recover)
- On FIFO queues when the exact message order must never be broken
Replay (redrive) discipline
Fix the root cause before redriving. Reprocessing a poison message without fixing the underlying bug sends it straight back to the DLQ — a pattern described as "re-poisoning" (Codemia, 2024–2026 practitioner writing).
Best practice for DLQ replay (Codemia; designgurus.substack.com):
- Diagnose the failure class (poison pill vs transient)
- Fix the consumer code or data issue first
- Redrive deliberately — treat replay as an operator-triggered action with a log of what was replayed and when
- Scope the redrive by failure cause, consumer code version, and idempotency guarantees — do not replay everything at once
Because redriven messages may be delivered twice (once when they originally failed, once on replay), idempotent consumers are a prerequisite for safe redrive (designgurus.substack.com). See Idempotency.
DLQ and related patterns
| Pattern | Relationship to DLQ |
|---|---|
| Idempotency | Prerequisite for safe redrive — idempotent consumers ensure replayed messages don't create duplicate orders, charges, or state changes |
| Circuit Breaker | Upstream complement — circuit breaker stops traffic to a failing service; DLQ catches messages that slip through before the circuit opens |
| Outbox Pattern | Upstream complement — Outbox guarantees at-least-once delivery to the queue; DLQ handles failures once the message reaches the consumer |
| Saga Pattern | DLQs surface saga step failures that need compensating transactions |
| Eventual Consistency | DLQs are part of the recovery path that makes eventual consistency achievable under failure |
Key terms
| Term | Meaning |
|---|---|
| Poison pill / poison message | A message that deterministically fails processing regardless of retry count |
| Redrive policy | Rules (chiefly maxReceiveCount) governing when a message moves to the DLQ |
| maxReceiveCount | How many times a message can be received before being moved to the DLQ |
| TTL | Time-to-live; a message that expires before processing is moved to the DLQ |
| Dead letter topic (DLT) | Kafka-native equivalent of a DLQ — a regular topic used as the DLQ terminal state |
| Non-blocking retry | Retry pattern where failed messages are forwarded to separate retry topics, freeing the main consumer thread |
| Redrive-to-source | AWS SQS feature to move DLQ messages back to the original queue for reprocessing |
Benchmarks / operational data (as-of 2026-08-21)
| Data point | Value | Source |
|---|---|---|
| Messages accumulated in unowned DLQ over 14 months | 2.3 million | syarif, Level Up Coding, 2026-08-19 |
| Order events lost in Black Friday incident (pre-DLQ) | 6,200 | syarif, Level Up Coding, 2026-08-19 |
| Duration of inventory outage triggering the DLQ addition | 40 minutes | syarif, Level Up Coding, 2026-08-19 |
| Bad payment message retry count before DLQ isolation (cited case) | >10,000 retries over 6 hours | designgurus.substack.com, 2025–2026 |
What practitioners report
The dominant practitioner concern in 2026 is not whether to use a DLQ, but governance after adding one: ownership, retention policy, PII review, and runbook creation are frequently absent at the time of creation. The syarif case (Level Up Coding, August 2026) is the most detailed first-person account of this failure mode in the current literature.
A secondary concern is replay safety: without idempotent consumers and a fix-first discipline, redrive operations re-poison the queue and compound the original failure.