On this page
Two-Phase Commit (2PC)
Two-Phase Commit (2PC)
Two-Phase Commit (2PC) is a distributed atomic commitment protocol that ensures all nodes in a distributed system either commit or roll back a transaction together. One node acts as the coordinator (transaction manager); the others are participants (resource managers). It enforces ACID guarantees — specifically atomicity — across multiple independent databases or services.
How it works
Phase 1 — Prepare (Voting):
The coordinator sends a PREPARE message to all participants. Each participant executes the operation locally (but does not commit), writes to a write-ahead log for durability, acquires all necessary locks, and votes YES or NO. A YES vote is a binding promise: the participant cannot unilaterally abort after voting YES. (Hossein Nejati Javaremi, Medium, April 2, 2026)
Phase 2 — Commit or Abort:
If all participants vote YES, the coordinator sends COMMIT to all; participants commit and release locks. If any participant votes NO, the coordinator sends ABORT to all; participants roll back. (Hossein Nejati Javaremi, Medium, April 2, 2026)
PostgreSQL implements 2PC natively via PREPARE TRANSACTION, COMMIT PREPARED, and ROLLBACK PREPARED. After PREPARE TRANSACTION, "there is a very high probability that it can be committed successfully, even if a database crash occurs before the commit is requested." COMMIT PREPARED "can be issued from any session, not only the one that executed the original transaction." (PostgreSQL 18 docs, August 13, 2026)
PostgreSQL's own docs caution: "PREPARE TRANSACTION is not intended for use in applications or interactive sessions. Its purpose is to allow an external transaction manager to perform atomic global transactions across multiple databases or other transactional resources. Unless you're writing a transaction manager, you probably shouldn't be using PREPARE TRANSACTION." (PostgreSQL 18 docs, August 13, 2026)
XA standard
The X/Open XA (eXtended Architecture) standard is the database-level specification for 2PC across heterogeneous data sources. XA roles: RM (Resource Manager), TM (Transaction Manager / coordinator), AP (Application Program). In a complete global transaction, TM interacts with RMs eight times — contributing to coordination complexity and performance overhead. (Alibaba Cloud, Zhu Jinjun, October 21, 2021)
PostgreSQL "follows the features and model proposed by the X/Open XA standard, but does not implement some less often used aspects." (PostgreSQL 18 docs, August 13, 2026)
Almost all major databases support XA: MySQL (InnoDB, since 5.0), PostgreSQL, Oracle, SQL Server. Java's JTA (Java Transaction API) and Spring expose XA to application code. (Alibaba Cloud, October 2021)
CAP theorem positioning
2PC is a CP (Consistency + Partition-tolerance) protocol in the CAP theorem framing. During a network partition, if the coordinator cannot reach a participant, the transaction cannot complete — the system blocks rather than proceeding with uncertainty. This makes it unsuitable for systems prioritising availability. (Hossein Nejati Javaremi, Medium, April 2, 2026)
Ecommerce applications
The canonical ecommerce 2PC scenario is an order creation touching multiple independent services simultaneously: Payment Service (charge the card), Inventory Service (deduct stock), Order Service (create the order record). Without coordination, partial failures leave inconsistent states: inventory deducted but payment failed, or payment charged but no order created. (Red Hat Developer, Keyang Xiang, October 2018, updated August 2023)
TCC (Try-Confirm-Cancel) is a 2PC variant designed specifically to release resource locks faster by using intermediate "reserved" states instead of database locks. It is the variant most commonly used in ecommerce-scale distributed transaction frameworks:
- Order system: Try = create "pending payment" order; Confirm = set "completed"; Cancel = set "canceled"
- Payment system: Try = freeze RMB 100 (balance unchanged); Confirm = debit RMB 100; Cancel = unfreeze
- Inventory system: Try = freeze 1 unit (stock count unchanged); Confirm = decrement stock; Cancel = unfreeze
- Membership/credits system: Try = prepare 10 credits (balance unchanged); Confirm = add credits; Cancel = clear preparation
(Alibaba Cloud, Zhu Jinjun, October 21, 2021 — Taobao book purchase example)
Seata is the dominant open-source distributed transaction framework for this pattern in the Java/microservices ecosystem. Originated as TXC (Taobao internal, 2014) → GTS (Alibaba Cloud, 2016) → Seata (open-source, 2019). Seata's XA mode uses a data source proxy to avoid business-code intrusion. Its AT mode automatically creates before/after snapshots for rollback without business code changes. (Alibaba Cloud, October 2021)
eBay's GRIT protocol is cited as evidence that 2PC-like guarantees are needed at enterprise ecommerce scale, but standard 2PC doesn't fit — eBay built a custom distributed transaction protocol. (Red Hat Developer, Bilgin Ibryam, September 2021)
Kafka transactional producer uses a two-phase protocol internally for exactly-once message delivery within a single Kafka cluster. (Hossein Nejati Javaremi, Medium, April 2, 2026)
Where 2PC survives in 2026
- Distributed SQL databases — CockroachDB and Google Spanner use 2PC internally at the storage layer, within a single team's ownership boundary. The application gets distributed ACID transactions without managing the protocol itself. (Java Code Geeks, Eleftheria Drosopoulou, July 28, 2026)
- Kafka transactional producer — comparable two-phase handshake to publish a message batch atomically. (Java Code Geeks, July 28, 2026)
- Legacy XA integration — integrating with third-party black-box or legacy systems that implement the 2PC/XA specification. (Red Hat Developer, September 2021)
- Financial/ledger systems with strict atomicity requirements and modest throughput needs. (Hossein Nejati Javaremi, Medium, April 2, 2026)
Failure modes
1. The blocking problem (most critical) If the coordinator crashes after all participants have voted YES but before sending the Phase 2 decision, participants are stuck in the "prepared" state — holding locks indefinitely until the coordinator recovers. They cannot commit (others may have received ABORT) and cannot abort (they promised to commit). PostgreSQL docs: "It is unwise to leave transactions in the prepared state for a long time. This will interfere with the ability of VACUUM to reclaim storage, and in extreme cases could cause the database to shut down to prevent transaction ID wraparound." (PostgreSQL 18 docs, August 13, 2026)
Full TM failure taxonomy (7 scenarios, Alibaba Cloud engineering):
- TM crash before phase 1 queries → no action on recovery
- TM crash after querying some RMs → send rollback to those who received query
- TM crash after querying all RMs but before logging phase 1 complete → rollback all
- TM crash after logging phase 1 complete → commit or rollback based on log
- TM crash before commit/abort log in phase 2 → based on log after recovery
- TM crash before logging phase 2 complete → re-issue decision based on log
- TM crash after logging phase 2 complete → no action needed (Alibaba Cloud, Zhu Jinjun, October 21, 2021)
2. Cloggage When 2PC holds locks during the commit phase, unrelated transactions that need the same data are also blocked. Under high contention this cascades across the system. This is distinct from the coordinator-crash blocking problem and persists even in healthy operation. (abadid, Calvin/FaunaDB researcher, Hacker News, January 2019)
3. Split-brain If the coordinator crashes after sending COMMIT to some but not all participants, the system ends up with some nodes committed and others not. Requires manual DBA intervention. (Alibaba Cloud, October 2021)
4. Deadlock risk Two concurrent 2PC transactions can mutually lock each other when each holds a resource the other needs. (Alibaba Cloud, October 2021)
5. Slowest-participant degradation Overall throughput is bounded by the slowest participant's response time. One degraded service brings the whole transaction down.
6. Monitoring gap Practitioners often don't monitor for 2PC failures (coordinator crashes, stuck transactions, lock timeouts). When failures occur they are discovered from user complaints, not alerting. (exabrial, Hacker News, January 2019)
3PC (Three-Phase Commit) as a partial fix:
3PC adds a PRE-COMMIT phase to reduce the blocking problem by allowing participants to infer the coordinator's intent if it crashes. However, it introduces consistency problems: if the coordinator crashes after sending PRE-COMMIT, some nodes auto-commit on timeout even if the coordinator intended abort. Alibaba Cloud engineering team: "3PC improves availability at the cost of reduced consistency... which is unacceptable for many OLTP systems." 3PC is "rarely applied to practical cases" — teams moved to Saga Pattern instead. (Alibaba Cloud, October 2021; YouTube practitioner source, July 2025)
Comparison with alternatives
| Dimension | 2PC / XA | TCC | Saga Pattern | Outbox Pattern |
|---|---|---|---|---|
| Consistency model | Strong (ACID atomicity) | Strong (via reserved states) | Eventual | Eventual (local atomicity) |
| Lock duration | Held across all phases | Short (Try phase only) | No cross-service locks | No cross-service locks |
| Transaction type | Short-lived only | Short-to-medium | Long-lived | Long-lived |
| Scalability | Low | Medium | High | High |
| Coordinator | Required, SPOF | Required | Optional (orchestration) | None |
| Isolation | Full | Partial (intermediate state visible) | None by default | None by default |
| Implementation cost | Low (XA proxy) | High (explicit Try/Confirm/Cancel per service) | High (compensating transactions) | Medium |
| Cross-vendor fit | Poor (requires XA support) | Medium | Good | Good |
| Throughput cost | ~30–40% reduction (as-of 2026-07-28) | Lower than XA | Minimal | Minimal |
(Baeldung, June 2024; Red Hat Developer, September 2021; Java Code Geeks, July 28, 2026)
Microservices.io (Chris Richardson): "2PC is not an option" for microservices with the database-per-service pattern. The site explicitly positions Saga as the replacement for 2PC, linking to a diagram titled "From_2PC_To_Saga.png". (microservices.io, copyright 2026)
Red Hat Developer (Bilgin Ibryam): 2PC is warranted in four specific scenarios: (1) writes to disparate resources cannot be eventually consistent, (2) heterogeneous data sources, (3) exactly-once message processing required and services cannot be made idempotent, (4) integrating with third-party black-box or legacy systems implementing 2PC. (Red Hat Developer, September 2021, updated October 2023)
Java Code Geeks (July 2026): "'2PC didn't get rejected on paper. It got avoided in practice, one incident at a time.'" (Java Code Geeks, Eleftheria Drosopoulou, July 28, 2026)
Practitioner decision questions
From Java Code Geeks (July 2026):
- Does one team own every participant, or does the transaction cross ownership and vendor boundaries?
- Can the business tolerate a brief window of inconsistency, or does a regulator or ledger require otherwise?
- Is there a realistic compensating action for every step, including hard-to-undo ones (sending email, charging card)?
- What happens to a saga that fails halfway and never receives its compensating event — is there a timeout and reconciliation job?
- Would an optimized, single-vendor 2PC implementation actually avoid the blocking problems, or is that assumption untested?
Practitioner advice: "make the surface area in 2PC as small as possible to minimize impact of a failed 2PC... you are probably not monitoring for failures. Start doing that." (exabrial, Hacker News, January 2019)
Practitioner alternative: "make all operations idempotent and ensure they are all run at least once" — Idempotency as the practical replacement pattern. (dilatedmind, Hacker News, January 2022)
VoltDB edge case: deterministic execution eliminates the need for 2PC entirely by ordering stored procedure invocations serially. (arielweisberg, third VoltDB engineer, Hacker News, January 2019)
Contradictions
2PC is categorically unsuitable for microservices (microservices.io) vs. warranted in specific cases (Red Hat). Chris Richardson (microservices.io, 2026): "2PC is not an option" for database-per-service microservices architecture. Bilgin Ibryam (Red Hat Developer, September 2021): 2PC is a valid choice in four scenarios: writes can't be eventually consistent, heterogeneous data sources, exactly-once required, legacy/third-party XA integration. Note: This is a framing difference, not a factual contradiction. Richardson optimises for modern cloud-native defaults; Ibryam advises from consulting realism on migration/legacy contexts.
Indefinite blocking on coordinator crash (textbook) vs. 1.5–5 seconds in production (NDB practitioner). Textbook framing: if coordinator crashes after phase 1, participants hold locks indefinitely. jamesblonde (NDB/MySQL Cluster practitioner, Hacker News, January 2019): "if a participant fails, yes you have to wait until a failure detector indicates it has failed. But in production systems, that is typically 1.5 to 5 seconds." He argues production-grade failure detectors make "indefinite blocking" an overstatement for well-implemented systems.
Banking as canonical 2PC use case vs. real banking uses deferred settlement. Many textbooks and tutorials cite bank transfers as the canonical 2PC example. Hacker News community (2012 Starbucks thread): Real networked credit banking does not use 2PC. ATMs use deferred settlement and reconciliation. "Networked credit banking was basically worked out in 4th century BC Egypt." Real financial systems prefer idempotent operations with eventual reconciliation over distributed locking.
3PC solves 2PC's blocking problem vs. 3PC introduces new consistency problems. Standard framing: 3PC adds PRE-COMMIT to reduce blocking by allowing participants to resolve without coordinator. Alibaba Cloud engineering (October 2021): "3PC improves availability at the cost of reduced consistency... which is unacceptable for many OLTP systems." Under network partition, nodes may auto-commit on timeout even if coordinator intended abort.
Atomikos (contrarian): 2PC was discarded for organizational reasons, not technical ones. Mainstream position (microservices.io, Red Hat, Baeldung, JCG): 2PC is technically unsuitable for polyglot microservices because of blocking, SPOF, and cross-vendor XA complexity. Atomikos (cited in Java Code Geeks, July 2026): "within a single team's ownership boundary, an optimized 2PC implementation can avoid SPOF and blocking problems. Some teams adopted sagas less because 2PC was unworkable and more because they'd been told it was."
Key terms
| Term | Meaning |
|---|---|
| Coordinator / TM | Transaction Manager — sends PREPARE and COMMIT/ABORT; the single point of failure |
| Participant / RM | Resource Manager — holds locks, votes, and executes on coordinator's decision |
| XA | eXtended Architecture — X/Open standard for 2PC across heterogeneous data sources |
| Prepared state | State a participant enters after voting YES in phase 1; holds locks until phase 2 completes |
| TCC | Try-Confirm-Cancel — 2PC variant using reserved states instead of DB locks; dominant in ecommerce frameworks like Seata |
| 3PC | Three-Phase Commit — adds PRE-COMMIT phase to reduce blocking; rarely used in practice due to consistency trade-offs |
| Cloggage | Cascade blocking of unrelated transactions caused by 2PC lock-holding under high concurrency (practitioner term, abadid, HN 2019) |
| JTA | Java Transaction API — Java standard for managing XA transactions |
| Seata | Open-source distributed transaction framework from Alibaba; TXC→GTS→Seata (2014→2016→2019) |
| GRIT | eBay's custom distributed transaction protocol built because standard 2PC didn't fit their scale |