An online order touches stock, payments and shipping, and in a service architecture each of those lives in its own database. There is no single transaction that spans them. Two-phase commit could stitch them together, but it holds locks across network round trips and blocks when the coordinator disappears, and most payment providers and SaaS APIs will never take part in it anyway. The saga pattern is the practical answer: split the business transaction into a sequence of local transactions, each committed on its own, and pair every step that might need undoing with a compensating action that semantically reverses it.
The idea comes from Hector Garcia-Molina and Kenneth Salem's 1987 paper on long-lived transactions. It is also widely misapplied: a saga provides no isolation, and its correctness rests on details that are easy to skip, namely idempotent steps, a durable record of progress, and a careful choice of the point of no return. This article builds a saga from first principles, walks through an order checkout, and gives you an orchestrator and a list of production failure modes.
What a saga guarantees, and what it does not
A saga guarantees that either every step eventually completes, or every completed compensatable step is eventually compensated. That is atomicity in outcome, reached over time, not at an instant. In between, other transactions can see partial results: stock that has been reserved for an order that will be cancelled, or a card authorisation that is about to be voided.
In ACID terms, sagas keep atomicity (eventually), consistency (if every step and compensation preserves your invariants) and durability (each local commit is durable), and drop isolation completely. Everything that makes sagas hard follows from that missing I. If you need a real atomic commit across stores you own, read two-phase commit first, and understand what blocking costs before rejecting it.
Compensation is semantic, not physical. You cannot un-send an email or un-charge a card by rolling back a row. A compensation is a new forward action, such as a refund, a release or an apology message, that returns the business to an acceptable state. Some effects have no compensation at all, and the pattern gives you a place to put them.
Three kinds of step
Classify every step before you write code. The classification, not the framework, decides whether your saga is correct.
- Compensatable steps can be undone by a compensating transaction. Reserving stock (undo: release it) and authorising a card (undo: void the authorisation) are typical.
- The pivot step is the go/no-go decision. Once it commits, the saga will run to completion. It is either the last compensatable step or the first step that cannot be compensated. Capturing payment is a common pivot: after money moves, you finish the order or refund it through a separate business process.
- Retriable steps come after the pivot and are guaranteed to succeed if retried enough times, because nothing about the business can make them fail permanently. Creating a shipment request and sending a confirmation email belong here.
Order the steps so that the steps most likely to fail run first and cheap-to-undo steps come before expensive ones. A saga that captures payment before checking stock will refund customers daily; one that checks stock first will rarely touch the refund path.
Orchestration or choreography
There are two ways to drive the sequence. In choreography, each service listens for the previous service's event and publishes its own: StockReserved triggers the payment service, PaymentAuthorised triggers capture, and so on. In orchestration, a single orchestrator holds the saga's state machine and sends commands to each participant, waiting for replies.
| Concern | Choreography | Orchestration |
|---|---|---|
| Where the flow lives | Spread across services' event handlers | One state machine, in one place |
| Adding a step | Change several services and their subscriptions | Change the orchestrator definition |
| Visibility | Reconstructed from traces and logs | Query the saga log |
| Coupling | Services know each other's events | Participants know only their own commands |
| Good fit | Two or three steps, stable flow | Four or more steps, branching, timeouts, human steps |
Choreography gets hard to reason about at around four steps, because nobody can say which orders are stuck at which step without joining logs across teams. This article uses orchestration. Workflow engines such as Temporal, AWS Step Functions and Camunda implement the orchestrator for you, but you still owe them idempotent participants.
Worked example: order checkout
Take an order for two units of SKU 42, paid by card. The saga definition is a list of steps with their compensations:
| Step | Action (local transaction) | Compensation | Class |
|---|---|---|---|
| T1 | Inventory: reserve 2 x SKU 42 for order 9001 | Release reservation | Compensatable |
| T2 | Payments: authorise 59.90 on the card | Void authorisation | Compensatable |
| T3 | Payments: capture the authorisation | None within this saga (refund is a separate process) | Pivot |
| T4 | Shipping: create shipment | None; retried until it succeeds | Retriable |
| T5 | Notifications: send confirmation | None; retried until it succeeds | Retriable |
Happy path: T1 to T5 commit and the saga ends COMPLETED. Now suppose the card issuer declines at T2. The orchestrator records T2 as FAILED, runs C1 to release the stock, and ends COMPENSATED. The customer sees a declined payment and the stock returns to sale. Suppose instead T3 times out. The orchestrator does not know whether the capture happened, so it must not compensate yet: it retries T3 with the same idempotency key until it gets a definite answer. If the provider says captured, the saga continues forward; if it says the authorisation expired, C2 has nothing left to void, so the saga records it as done and compensates C1. A timeout is never a failure; it is an unknown.
A durable orchestrator
The orchestrator's only real job is to never lose its place. It records each transition in a saga log in its own database before it sends the next command, and on restart it resumes every saga that is not in a terminal state. The sketch below is deliberately small; the parts that matter are the ordering of writes and the idempotency keys.
import uuid
class Step:
def __init__(self, name, action, compensate=None, kind="compensatable"):
self.name, self.action, self.compensate, self.kind = name, action, compensate, kind
class Unknown(Exception): # timeout or ambiguous reply: outcome not known
pass
class Failed(Exception): # definite business failure: safe to compensate
pass
def run_saga(db, saga_id, steps, ctx):
state = db.load_saga(saga_id) or db.create_saga(saga_id, ctx)
i = state.next_step
while i < len(steps) and state.status == "RUNNING":
step = steps[i]
key = f"{saga_id}:{step.name}" # same key on every retry
try:
result = step.action(ctx, idempotency_key=key)
ctx.update(result or {})
db.record(saga_id, step=i, status="DONE", ctx=ctx) # commit before moving on
i += 1
except Unknown:
db.schedule_retry(saga_id, step=i) # resume later, do NOT compensate
return "PENDING"
except Failed:
if i > pivot_index(steps):
db.schedule_retry(saga_id, step=i) # past the pivot: only forward
return "PENDING"
db.record(saga_id, step=i, status="FAILED", ctx=ctx)
return compensate(db, saga_id, steps[:i], ctx)
db.finish(saga_id, "COMPLETED")
return "COMPLETED"
def compensate(db, saga_id, done_steps, ctx):
for i in reversed(range(len(done_steps))):
step = done_steps[i]
if step.compensate:
key = f"{saga_id}:{step.name}:undo"
step.compensate(ctx, idempotency_key=key) # must be retriable until it succeeds
db.record(saga_id, step=i, status="COMPENSATED", ctx=ctx)
db.finish(saga_id, "COMPENSATED")
return "COMPENSATED"
def pivot_index(steps):
return next(i for i, s in enumerate(steps) if s.kind == "pivot")Three rules are carried by this code. The log write happens before the next step, so a crash between them replays at most one step, and idempotency makes the replay harmless. A timeout parks the saga for retry instead of compensating, because compensating an action that did happen is how customers get charged for cancelled orders. And compensations are themselves retriable: a compensation that fails permanently is a bug that needs a human, so it must page someone rather than be silently dropped.
When the orchestrator and participants talk through a message broker, sending the command and recording the transition must not be a dual write. Put the outgoing command in an outbox table in the same local transaction as the saga log update, and relay it from there, as described in transactional outbox. Participants do the same for their replies.
Idempotency is not optional
Messages are delivered at least once, orchestrators replay after crashes, and retries follow timeouts. Every participant must therefore turn a repeated command into a no-op that returns the original result. The standard implementation is a processed-requests table keyed by idempotency key, written in the same local transaction as the business change:
BEGIN;
INSERT INTO processed_requests (idempotency_key, response)
VALUES ('9001:reserve', NULL)
ON CONFLICT (idempotency_key) DO NOTHING
RETURNING idempotency_key; -- no row returned means: already handled
-- only if a row came back:
UPDATE stock SET reserved = reserved + 2 WHERE sku = 42 AND available - reserved >= 2;
UPDATE processed_requests SET response = '{"reserved":2}' WHERE idempotency_key = '9001:reserve';
COMMIT;Two subtleties matter. A compensation can arrive before the action it undoes, when messages are reordered; record the compensation's key anyway and make the late action check for it and refuse to apply. And external APIs need the same treatment: many payment providers accept an idempotency key, and where one does not, query for the effect before retrying.
Living without isolation
Because each step commits immediately, concurrent sagas and ordinary requests see intermediate state. The classic anomalies appear in business form: a lost update when two sagas adjust the same account, a dirty read when a report counts stock reserved for an order that is about to be cancelled, and a fuzzy read when a saga checks a credit limit that another saga changes before the pivot. The saga literature names a set of countermeasures, and most real systems combine several:
- Semantic lock. Mark records a saga is working on with a status such as PENDING_APPROVAL. Other requests either wait, fail fast, or treat the record as provisional. The compensation or final step clears the flag. This is an application-level lock, so it needs a timeout and an owner.
- Commutative updates. Design operations so their order does not matter: debit and credit rather than set-balance. Commutative steps and compensations cannot lose each other's updates.
- Pessimistic ordering. Reorder steps so the risky ones come after the pivot, or so the step that others read is updated last.
- Reread value. Before overwriting, reread and verify the record has not changed since the saga read it, an optimistic check on a version column. If it has, abort or restart the saga.
Semantic locks held by crashed orchestrators need the same care as any lease: expire them, and fence late writers with a monotonically increasing token, as in fencing tokens.
Failure modes you will meet
- Compensating on timeout. The most expensive saga bug. The action happened; the compensation undoes it; a later retry of the action does it again. Treat timeouts as unknown and resolve them before deciding direction.
- Non-idempotent participant. A redelivered command reserves stock twice or sends two emails. Fix with a processed-requests table and a key derived from saga ID and step name, never a fresh UUID per attempt.
- Compensation that can fail. Voiding an authorisation after it expired, or releasing stock that another process already shipped. Compensations need their own retry policy, an alert when they exhaust it, and a manual resolution queue.
- Lost orchestrator state. Keeping progress in memory, or in a cache, means a restart orphans every in-flight saga. The saga log must live in durable storage with the same backups as the business data.
- Definition changes mid-flight. Deploying a new step order while sagas are running sends old sagas down new paths. Version saga definitions and let in-flight sagas finish on the version they started with.
- Retry storms against a failing participant. Thousands of parked sagas retry at once when a dependency recovers. Use jittered backoff and a circuit breaker per participant.
Operating sagas in production
Treat the saga log as your primary dashboard: count by status and step, age of the oldest RUNNING saga per step, compensation rate per step, and sagas parked in manual resolution. Alerting on the oldest-running age catches stuck sagas long before customers do. Propagate the saga ID as a trace attribute and in every log line so one query reconstructs an order's history across services.
Set per-step deadlines as well as retry limits; how long a reservation may stay PENDING is a business decision, so agree it with the product owner. Test compensations with a synthetic order in staging that fails at each step in turn and must end COMPENSATED with stock and money restored.
Trade-offs and when not to use a saga
Sagas buy availability and service autonomy with complexity and weaker guarantees. They fit when steps span services or external providers and the business tolerates brief intermediate states. They fit poorly when an invariant must never be visible as violated, such as a ledger that must always balance at every instant; keep that invariant inside one database transaction and move the saga boundary around it. If two steps share a database, merge them into one local transaction: the cheapest saga step is the one you did not need.
What to do next
- Write your workflow as a table of steps, compensations and classes, and mark exactly one pivot.
- Reorder steps so likely failures and cheap undos come first and irreversible effects come after the pivot.
- Give every participant a processed-requests table and derive idempotency keys from saga ID plus step name.
- Store saga progress in a durable log and send commands through an outbox in the same transaction.
- Treat timeouts as unknown: retry with the same key, and only compensate on a definite failure.
- Choose isolation countermeasures per shared record: semantic lock, commutative update or reread check.
- Alert on the age of the oldest running saga per step and on compensation failures, and test every failure point in staging.