On this page
concept

Dead Letter Queue (DLQ)

Created 2026-08-21 25 connections

Dead Letter Queue (DLQ)

A dead letter queue (DLQ) is a special message queue that stores messages a system cannot process after exhausting its retry budget. In ecommerce, DLQs are the last line of defence in event-driven order pipelines: when a payment webhook, inventory update, or fulfilment event fails processing, the DLQ catches it rather than letting it be silently dropped or cause a consumer to block indefinitely.

How it works

A DLQ operates alongside a regular message queue using a redrive policy — a set of rules (chiefly maxReceiveCount) that govern when a message is moved out of the main queue into the dead letter queue. The message's failure reason is typically one of (AWS "What is a DLQ?" documentation, updated 2026-06-08):

  • Erroneous message content — hardware, software, or network corruption changes the payload; receiver rejects it
  • Changes in receiver's system — consumer API schema changes or the entity the message references no longer exists
  • TTL expiry — the message's time-to-live expires before the consumer can process it
  • Message size limit exceeded — payload exceeds the queue's maximum message size

Messages remain in the DLQ until developers inspect, fix the root cause, and redrive them back to the source queue (AWS docs, 2026-06-08).

Poison messages vs transient failures

Two distinct failure classes determine the correct DLQ handling strategy (Redpanda / Dunith Danushka, 2022; widely cited in 2024–2026 practitioner writing):

Failure classCharacteristicsDLQ handling
Non-transient (poison pill)Deterministic — always fails regardless of how many retries. Caused by malformed JSON, schema mismatches, validation failures, consumer code bugsRoute directly to DLQ; do not retry — retrying wastes resources and never succeeds
TransientNon-deterministic — may self-heal. Caused by temporary downstream unavailability, network timeouts, database locksRetry with exponential backoff first; only park in DLQ if all retry attempts are exhausted

The Redpanda source (Dunith Danushka) was published November 2022. The taxonomy of poison-pill vs transient failures remains standard in 2026 practitioner writing, but Spring Kafka API specifics (ErrorHandlingDeserializer versions) may have changed. Treat code examples as illustrative, not current.

Retry patterns before the DLQ

Practitioners distinguish between two retry approaches before a message reaches the DLQ (Redpanda blog, 2022; Codemia knowledge hub):

Blocking (synchronous) retry

Consumer thread is suspended while the same message is retried. Drawback: clogs the consumer thread; prevents offset commit; all downstream messages from the same partition are held behind the failing one (Redpanda, 2022).

Non-blocking (asynchronous) retry with backoff

Failed messages are forwarded to one or more retry topics (separate from the DLQ), processed by a dedicated retry consumer. This frees the main consumer thread to continue with fresh messages.

A common Spring Kafka pattern (Redpanda, 2022) uses @RetryableTopic with exponential backoff across multiple retry topics before the message finally lands in the DLQ:

orders → orders-retry-0 (1s delay) → orders-retry-1 (2s delay) → orders-retry-2 (4s delay) → orders-dlt

In Kafka-based systems, the DLQ is typically a dead letter topic (DLT) — a regular topic used as the DLQ terminal state (Redpanda, 2022).

DLQ in ecommerce: key failure scenarios

Order event processing

Amazon's order processing uses DLQs to quarantine corrupted events that would otherwise trigger cascading failures across inventory, payment, and shipping services (referenced in designgurus.substack.com, 2025–2026 writeup).

Payment processing

Without a DLQ, a single malformed payment event in a Kafka partition blocks all subsequent payments from that partition. A cited case: a consumer retried a bad message >10,000 times over six hours, blocking every payment from that partition — DLQ would have isolated the bad message immediately (designgurus.substack.com, 2025–2026).

Inventory and fulfilment events

A Black Friday incident described by syarif (Level Up Coding, August 2026): a downstream inventory service went down for 40 minutes; 6,200 order events failed with no retry path. Post-incident, a Kafka DLQ was added to capture failed events rather than dropping them.

The governance problem

The most detailed operational account in 2026 practitioner writing (syarif / Level Up Coding, August 2026) documents a DLQ that became a liability:

  • Kafka DLQ (orders.dlq) was added after Black Friday incident with a three-line PR description: "adds DLQ for resilience."
  • No runbook, no owner, no retention policy, no scheduled review was assigned at creation
  • 14 months later: 2.3 million messages had accumulated, some going back 14 months
  • A compliance audit found the DLQ contained customer email addresses, shipping details, and partial payment metadata — data categories the team's retention policy (71 days) applied to orders.created but had never been applied to orders.dlq

This case surfaces a structural risk: a DLQ is a data store, not just an error buffer. It inherits the PII and retention obligations of the messages it receives, but is rarely treated as such at creation time.

Source: syarif, "Before You Add a Dead Letter Queue," Level Up Coding / Medium, 2026-08-19

Configuration parameters

AWS SQS (as-of 2026-06-08)

  • maxReceiveCount — number of times a message is received before being sent to the DLQ; optimize to avoid sending transiently-failing messages to DLQ too early (AWS docs)
  • Redrive policy — rules defining source queue ↔ DLQ interaction; includes maxReceiveCount
  • Redrive-to-source — SQS feature to replay DLQ messages back to their original queue (or a different destination) on demand (AWS docs)
  • For FIFO queues: the DLQ must also be a FIFO queue to preserve message ordering (AWS docs)

When NOT to use a DLQ (AWS docs, 2026-06-08)

  • When messages must be retried indefinitely (e.g., waiting for a dependent service to recover)
  • On FIFO queues when the exact message order must never be broken

Replay (redrive) discipline

Fix the root cause before redriving. Reprocessing a poison message without fixing the underlying bug sends it straight back to the DLQ — a pattern described as "re-poisoning" (Codemia, 2024–2026 practitioner writing).

Best practice for DLQ replay (Codemia; designgurus.substack.com):

  1. Diagnose the failure class (poison pill vs transient)
  2. Fix the consumer code or data issue first
  3. Redrive deliberately — treat replay as an operator-triggered action with a log of what was replayed and when
  4. Scope the redrive by failure cause, consumer code version, and idempotency guarantees — do not replay everything at once

Because redriven messages may be delivered twice (once when they originally failed, once on replay), idempotent consumers are a prerequisite for safe redrive (designgurus.substack.com). See Idempotency.

PatternRelationship to DLQ
IdempotencyPrerequisite for safe redrive — idempotent consumers ensure replayed messages don't create duplicate orders, charges, or state changes
Circuit BreakerUpstream complement — circuit breaker stops traffic to a failing service; DLQ catches messages that slip through before the circuit opens
Outbox PatternUpstream complement — Outbox guarantees at-least-once delivery to the queue; DLQ handles failures once the message reaches the consumer
Saga PatternDLQs surface saga step failures that need compensating transactions
Eventual ConsistencyDLQs are part of the recovery path that makes eventual consistency achievable under failure

Key terms

TermMeaning
Poison pill / poison messageA message that deterministically fails processing regardless of retry count
Redrive policyRules (chiefly maxReceiveCount) governing when a message moves to the DLQ
maxReceiveCountHow many times a message can be received before being moved to the DLQ
TTLTime-to-live; a message that expires before processing is moved to the DLQ
Dead letter topic (DLT)Kafka-native equivalent of a DLQ — a regular topic used as the DLQ terminal state
Non-blocking retryRetry pattern where failed messages are forwarded to separate retry topics, freeing the main consumer thread
Redrive-to-sourceAWS SQS feature to move DLQ messages back to the original queue for reprocessing

Benchmarks / operational data (as-of 2026-08-21)

Data pointValueSource
Messages accumulated in unowned DLQ over 14 months2.3 millionsyarif, Level Up Coding, 2026-08-19
Order events lost in Black Friday incident (pre-DLQ)6,200syarif, Level Up Coding, 2026-08-19
Duration of inventory outage triggering the DLQ addition40 minutessyarif, Level Up Coding, 2026-08-19
Bad payment message retry count before DLQ isolation (cited case)>10,000 retries over 6 hoursdesigngurus.substack.com, 2025–2026

What practitioners report

The dominant practitioner concern in 2026 is not whether to use a DLQ, but governance after adding one: ownership, retention policy, PII review, and runbook creation are frequently absent at the time of creation. The syarif case (Level Up Coding, August 2026) is the most detailed first-person account of this failure mode in the current literature.

A secondary concern is replay safety: without idempotent consumers and a fix-first discipline, redrive operations re-poison the queue and compound the original failure.

Research agent · 2026-08-21