On this page
- Context and problem
- Definition
- Two implementation variants
- Choreography
- Orchestration
- Transaction types within a saga
- Ecommerce reference architecture
- Isolation: the missing ACID property
- Isolation countermeasures
- Compensating transactions
- Production failure modes
- Message broker guidance
- Orchestration tooling (as-of 2026)
- Monitoring
- 2PC vs Saga vs Eventual Consistency
- Key terms
Saga Pattern
Saga Pattern
A distributed transaction pattern that maintains data consistency across multiple microservices without using traditional two-phase commit (2PC). A saga breaks a cross-service business operation into a sequence of local transactions, each committed atomically within a single service. Failures trigger compensating transactions that undo the effects of preceding steps.
Context and problem
In a microservices architecture following the Database per Service pattern, each service owns its own database. A single business transaction — such as a checkout flow touching inventory, payment, order, and fulfilment services — cannot be wrapped in a single ACID database transaction. According to Chris Richardson's microservices.io pattern catalogue: "2PC is not an option." (microservices.io, ongoing — as-of 2026)
AWS Prescriptive Guidance characterises the problem as: the need for data integrity across multiple data stores, especially where the data store does not support 2PC, or where NoSQL databases without ACID transactions require multi-table updates. (AWS Prescriptive Guidance, current — as-of 2026)
Definition
From microservices.io (Chris Richardson): "A saga is a sequence of local transactions. Each local transaction updates the database and publishes a message or event to trigger the next local transaction in the saga. If a local transaction fails because it violates a business rule then the saga executes a series of compensating transactions that undo the changes that were made by the preceding local transactions." (microservices.io, as-of 2026)
Microsoft Azure Architecture Center definition: "A saga is a sequence of transactions that updates each service and publishes a message or event to trigger the next transaction step. If a step fails, the saga executes compensating transactions that counteract the preceding transactions." (Microsoft Azure, updated 2025-02-25)
Two implementation variants
Choreography
Services react to events published by other services — there is no central coordinator. Each service listens for domain events, executes a local transaction, and publishes outcome events.
Benefits (Azure Architecture Center, as-of 2025-02-25):
- No additional orchestrator service required
- No single point of failure
Drawbacks (Azure Architecture Center, as-of 2025-02-25):
- Workflow logic becomes hard to follow as the number of steps grows
- Risk of cyclic event dependencies between services
- Integration testing requires all participant services to be running
AWS Prescriptive Guidance characterises choreography as: "suitable for few participants, simpler to implement, no single point of failure but harder to track as participants grow." (AWS Prescriptive Guidance, as-of 2026)
Orchestration
A central orchestrator object tells each participant which local transaction to execute. The orchestrator owns the workflow state and issues commands; participants respond with success or failure.
Benefits (Azure Architecture Center, as-of 2025-02-25):
- Better suited for complex, multi-step workflows
- Avoids cyclic dependencies
- Provides clear separation of participant responsibilities from flow logic
- Centralised visibility for monitoring and retries
Drawbacks (Azure Architecture Center, as-of 2025-02-25):
- Design complexity — coordination logic must be built into the orchestrator
- The orchestrator can become a single point of failure if not made highly available
AWS names AWS Step Functions as its recommended orchestration engine for sagas, with built-in multi-AZ fault tolerance. (AWS Prescriptive Guidance, as-of 2026)
Google Cloud demonstrates orchestration via Google Cloud Workflows, using try/retry for transient failures and try/except to route to compensating steps on permanent failures. (Google Cloud Blog, 2022-02-05)
thecodeforge.io (Naren, practitioner, as-of 2026-07-27) reports: "In production, most teams start with choreography for low-risk flows (notification chains) and switch to orchestration for payment pipelines where every step must be auditable."
Choreography vs orchestration: which is simpler?
- thecodeforge.io (Naren, as-of 2026-07-27): "Choreography feels simpler at first. No single point of failure, no coordinator to crash. But the hidden cost is observability." Recommends orchestration for financial flows and cites a first-person anecdote: a four-service choreography saga took three weeks to debug; after moving to orchestration, the same bug took two hours to fix.
- oneuptime.com (Feb 2026) offers a service-count decision rule: 2–3 services → choreography; 4+ services needing central visibility → orchestration.
- microservices.io (Richardson): presents both variants without prescribing a default or service-count threshold.
Transaction types within a saga
Azure Architecture Center and thecodeforge.io both reference Chris Richardson's taxonomy from Microservices Patterns (Medium/@t.shalash1, Feb 2026):
- Compensable transactions — local transactions that can be undone by a corresponding compensating transaction
- Pivot transaction — the "point of no return"; if it succeeds, all subsequent steps are retriable. Can be the last compensable step or the first retriable step
- Retriable transactions — occur after the pivot; guaranteed to succeed eventually; must be idempotent
Ecommerce reference architecture
martinuke0 (practitioner blog, Jun 2026) provides a multi-service order saga for a checkout flow:
| Service | Responsibility | Data store | Compensation |
|---|---|---|---|
| Cart | Reserve items, provisional total | PostgreSQL | Delete reservation rows |
| Pricing | Apply discounts, tax calculation | Redis + Postgres | Revert price overrides |
| Payment | Capture funds via gateway | MySQL | Issue refund via gateway API |
| Inventory | Decrement stock, allocate location | CockroachDB | Increment stock, release allocation |
| Order | Persist final order, trigger fulfilment | DynamoDB | Delete order, publish cancellation |
(martinuke0 blog, Jun 2026 — reference architecture; specific data stores are illustrative)
Temporal blog (Jan 2025) describes the typical ecommerce saga sequence as: Order Creation → Payment Processing → Inventory Update → Shipping, with the rule: "If a payment fails, the system cancels the order and restores inventory."
Isolation: the missing ACID property
Sagas provide Atomicity, Consistency, and Durability (ACD) but not Isolation. Medium/@t.shalash1 (Feb 2026) states: "Because each local transaction commits individually before the entire saga is finished, the intermediate changes are immediately visible to other sagas and transactions."
Three data anomalies that result (Azure Architecture Center, as-of 2025-02-25; Medium/@t.shalash1, Feb 2026):
- Lost updates — one saga overwrites changes made by another saga without reading the intermediate update
- Dirty reads — a saga reads data being modified by another in-flight saga that later fails and rolls back
- Fuzzy/non-repeatable reads — two steps within the same saga read the same data and get different results because another saga modified it between reads
Isolation countermeasures
Six countermeasures enumerated by Azure Architecture Center (as-of 2025-02-25) and Medium/@t.shalash1 (Feb 2026), both attributing Chris Richardson's Microservices Patterns:
- Semantic lock — a compensable transaction sets a
*_PENDINGflag on records it modifies, blocking other sagas from acting on them until the flag is cleared. Trade-off: reduces concurrency. - Commutative updates — design updates as incremental operations that produce the same result regardless of order (e.g.
balance -= 20rather thanSET balance = 80). Trade-off: only works where operations can be expressed as reversible increments. - Pessimistic view — reorder saga steps to move "risky" updates (those granting rights or increasing resources) to the end as retriable transactions. Trade-off: risky operations are delayed, increasing perceived latency.
- Reread value — before updating, re-read the record to verify it has not changed since the last read; abort and restart if it has (a form of Optimistic Offline Lock). Trade-off: does not prevent dirty reads.
- Version file — keep a log of operations performed on a record so they can be logically reordered. Useful when compensation events arrive before the forward event due to network reordering.
- By value — select concurrency mechanism based on business risk: use sagas for low-risk operations, distributed transactions or extra safeguards for large, high-value transactions. Trade-off: adds complexity to decision logic.
Compensating transactions
From microservices.io (Richardson, as-of 2026): "a developer must design compensating transactions that explicitly undo changes made earlier in a saga rather than relying on the automatic rollback feature of ACID transactions."
oneuptime.com (Feb 2026) defines a compensation failure escalation path: Retry with backoff → if exhausted → Dead Letter Queue → alert operations team → manual resolution.
thecodeforge.io (Naren, as-of 2026-07-27) states: "Design compensations to be: Stateless — rely only on identifiers, not on transient context. Re-runnable — if a refund fails once, retrying should not double-charge. Auditable — log each compensation with a correlation ID."
Microsoft Azure Architecture Center (as-of 2025-02-25) explicitly notes: "Compensating transactions don't always work." Steps may not fail immediately but instead get blocked, requiring timeout mechanisms.
Production failure modes
thecodeforge.io (Naren, as-of 2026-07-27) documents the most common production incidents:
Compensation race (rollback clashes with slow retry): A payment step succeeds but the success notification to the orchestrator is delayed; the saga times out and starts compensating. Meanwhile, the original notification arrives after compensation has completed. Fix: "Add a status flag per saga instance ('compensating'). Reject any success notification after the flag is set. Use Compare-and-Swap to set the flag atomically."
Timeout-induced double execution: A saga step times out, compensation starts, but the original step actually completed slowly. The compensation undoes a completed step. Fix: use a state machine with an explicit 'compensating' state that rejects forward completions.
Partial compensation: Compensation step A succeeds, step B fails. Only path: keep retrying B with exponential backoff, then escalate. A manual intervention playbook is required.
Fintech incident cited (no source named): A saga without idempotency keys caused a double withdrawal of $2M when a message broker delivered the same event after a network partition healed. Lesson from thecodeforge.io: "Idempotency is not optional in sagas — it's the safety net." (as-of 2026-07-27 — anecdote; not independently verified)
[!unverified] The $2M double-withdrawal anecdote is sourced from thecodeforge.io (practitioner blog) without a named institution or corroborating source. Filed for pattern illustration only.
Message broker guidance
martinuke0 (Jun 2026): "Kafka provides ordered, durable logs and supports exactly-once semantics when paired with idempotent producers. Its native transactional API lets you write a message and commit a local DB transaction atomically, which is essential for sagas."
"RabbitMQ can be used for low-latency command-style messages, but lacks built-in log compaction, making replay more complex." (martinuke0, Jun 2026)
Orchestration tooling (as-of 2026)
- Temporal — most-frequently cited across 2025–2026 sources (Temporal blog Jan 2025; thecodeforge.io Jul 2026; martinuke0 Jun 2026; YouTube talks 2024–2026)
- AWS Step Functions — named by AWS as their recommended orchestration engine for sagas; includes built-in multi-AZ fault tolerance (AWS Prescriptive Guidance, as-of 2026)
- Camunda — named alongside Temporal as an orchestration alternative (martinuke0, Jun 2026)
- Eventuate Tram Sagas — Chris Richardson's own open-source framework; reference implementations at eventuate-tram-sagas-examples-customers-and-orders (microservices.io, as-of 2026)
- Google Cloud Workflows — demonstrated for saga orchestration (Google Cloud Blog, 2022-02-05)
Note: commercetools and Fluent Commerce do not publish explicit Saga pattern documentation, though their event-driven OMS architectures are structurally compatible with saga choreography. (Registry agent research, 2026-08-19)
Monitoring
thecodeforge.io (Naren, as-of 2026-07-27) recommends per-step saga metrics: active saga count, completed count, failed count by step. Alert thresholds cited: compensation rate >2% may indicate upstream instability; saga latency >5s could affect checkout UX; Kafka consumer lag >5000 signals back-pressure. (as-of 2026-07-27)
2PC vs Saga vs Eventual Consistency
thecodeforge.io (as-of 2026-07-27) comparison:
| Concept | Coordination | Consistency | Failure handling | Use case |
|---|---|---|---|---|
| Saga Pattern | Choreography or Orchestration | Eventual | Compensating transactions | Order fulfilment pipeline |
| Two-Phase Commit (2PC) | Centralized coordinator | Strong (ACID) | Rollback via coordinator; can block | Tightly coupled financial settlement |
| Eventual Consistency (no saga) | None | Eventual (no compensations) | Manual intervention or TTL | User profile updates, non-critical data |
Google Cloud Blog (2022-02-05): Implementation guidance using Google Cloud Workflows predates current orchestration tooling landscape. The pattern description remains valid but tool-specific code examples may not reflect current SDK versions.
Key terms
| Term | Meaning |
|---|---|
| Compensating transaction | A local transaction that undoes the effect of a preceding saga step |
| Choreography | Saga variant where services react to events with no central coordinator |
| Orchestration | Saga variant where a central orchestrator directs all participants |
| Semantic lock | An application-level lock (*_PENDING flag) used as a countermeasure for isolation absence |
| Pivot transaction | The "point of no return" step in a saga; once committed, preceding steps no longer need compensating |
| Idempotency | Property ensuring a repeated operation produces the same result; mandatory for saga participants |
| Dead Letter Queue | Queue for saga steps that have exhausted retries and require manual intervention |
| Eventual Consistency | The consistency guarantee provided by sagas — all services converge to a consistent state, but not instantly |