Most write-ups of event-driven architecture stop at the pattern names: outbox, saga, idempotency, CQRS. An order system is where you discover how they fit together, because it touches money, physical stock and a waiting customer. Get the fit wrong and you oversell the last unit, charge a card twice, or strand an order in a state nobody owns.
This article designs one order system end to end: the order as a state machine, the events that move it, the checkout path through inventory and payment with code, and the failures that actually happen. The underlying mechanisms have their own articles on this site, linked where they appear; here the focus is the domain design that uses them.
Scope and a worked load
The system takes orders for a mid-sized online retailer. Customers place an order containing one or more stock-keeping units (SKUs); the system must hold stock, authorise payment, hand the order to the warehouse, capture payment on shipment, and let the customer see status at every step. Cancellation is allowed until the order ships.
Assume 1,500 orders per minute normally, 20,000 in the first minute of a flash sale, and about 12 events per order across its life. The flash peak is roughly 333 orders and 4,000 events per second, which any modern broker handles. The hard problems are not throughput but correctness under retries, ordering, and partial failure across services that each own a database.
Three requirements drive everything: no oversell beyond an agreed tolerance, no double charge ever, and no silent stuck order: every order is either progressing or visible to someone as stuck.
The order is a state machine first
Before choosing topics or services, write the order's states and the only legal transitions between them. Most production bugs in order systems are illegal transitions nobody forbade.
| State | Entered when | Legal next states |
|---|---|---|
| PENDING | order accepted and persisted | STOCK_HELD, REJECTED, CANCELLED |
| STOCK_HELD | all lines reserved | CONFIRMED, REJECTED, EXPIRED, CANCELLED |
| CONFIRMED | payment authorised | SHIPPED, CANCELLED, REJECTED |
| SHIPPED | carrier accepted the parcel | DELIVERED, RETURNED |
| REJECTED / EXPIRED / CANCELLED | stock, payment or customer said no | terminal |
Two choices follow. First, only the order service decides the order's state; other services own their facts (stock, payments, shipments) and report them as events. Second, each order carries a monotonically increasing version. Every state change increments it and every event the order service emits carries it, so consumers can tell a stale event from a fresh one without trusting delivery order.
Event contracts
Every event shares an envelope. The business payload sits in data; everything a consumer needs for deduplication, ordering and tracing sits outside it.
{
"event_id": "0192f1c4-6a0e-7c11-9d3a-5b2f8e41a7d0",
"event_type": "OrderPlaced",
"schema_version": 3,
"order_id": "ord_8KQ2V7",
"order_version": 1,
"occurred_at": "2026-09-30T08:05:12.418Z",
"correlation_id": "chk_51ZP",
"causation_id": "cmd_PlaceOrder_51ZP",
"producer": "order-service",
"data": {
"customer_id": "cus_1934",
"currency": "EUR",
"lines": [{"sku": "SKU-4411", "qty": 2, "unit_price_minor": 1999}],
"total_minor": 3998
}
}event_id is globally unique and is the deduplication key. order_id is the partition key. order_version lets a consumer discard an event older than what it has already applied. correlation_id ties every event of one checkout together for tracing; causation_id names the command or event that caused this one, which is what you need when reconstructing why an order did something odd. Money is carried in integer minor units with an explicit currency, never as floating point.
Carry the facts a consumer needs to act without calling back: fulfilment needs lines and address, notifications need customer and total. Callbacks reintroduce coupling and read the current state rather than the state at the time of the event. Evolve schemas additively, bump schema_version on any change, and keep a contract test in CI.
Architecture and the checkout data flow
- The client sends
POST /orderswith anIdempotency-Keyheader. A retried request with the same key returns the original response instead of creating a second order; the mechanics are in idempotency architecture. - The order service inserts the order in PENDING and, in the same database transaction, inserts
OrderPlacedand aReserveStockcommand into an outbox table. It returns 202 Accepted with the order id. No broker call happens inside the request; that dual write is the classic bug the transactional outbox exists to remove. - A relay, either change data capture on the outbox table or a poller, publishes the rows to the broker with
order_idas the message key. - Inventory reserves stock and replies
StockReservedorStockRejected. The order service moves to STOCK_HELD and emitsAuthorizePayment. - Payment authorises the card and replies. On
PaymentAuthorizedthe order becomes CONFIRMED and emitsCommitStock, which turns the hold into a sale only if it is still HELD; fulfilment is requested onStockCommitted. Capture happens later, onShipped.
Because the key is the order id, every event for one order lands on one partition and is consumed in the order it was published. That is the only ordering guarantee you get and the only one you need: nothing in this design depends on ordering across orders.
Orchestrate checkout, choreograph the rest
A saga can be choreographed, with each service reacting to the previous one's event, or orchestrated, with one component issuing commands and reacting to replies. The trade-offs in general are in saga pattern architecture. For orders the answer is usually both, split by purpose.
The checkout path is orchestrated by the order service, as the handler below shows. It has a strict sequence, compensations that depend on how far it got, and a timeout per step, and that logic belongs in one place you can read and test. Downstream reactions that do not affect the order's outcome, such as emails, loyalty points, analytics and search indexing, are choreographed: they subscribe to OrderConfirmed or OrderShipped and nobody waits for them.
# Order service: the checkout process manager. One handler per reply event.
def handle(event, db, outbox):
with db.transaction():
if db.exists("processed_events", event_id=event.event_id):
return # duplicate delivery: already applied
order = db.select_for_update("orders", order_id=event.order_id)
next_state, commands = TRANSITIONS.get((order.state, event.event_type), (None, []))
if next_state is None:
log.warning("ignored %s in state %s", event.event_type, order.state)
else:
order.state = next_state
order.version += 1
db.update("orders", order)
for cmd in commands: # e.g. AuthorizePayment, ReleaseStock
outbox.append(cmd.for_order(order)) # published later by the relay
db.insert("processed_events", event_id=event.event_id)
TRANSITIONS = {
("PENDING", "StockReserved"): ("STOCK_HELD", [AuthorizePayment]),
("PENDING", "StockRejected"): ("REJECTED", [NotifyCustomer]),
("PENDING", "CancelRequested"): ("CANCELLED", [ReleaseStock]), # no-op if none held
("STOCK_HELD", "PaymentAuthorized"): ("CONFIRMED", [CommitStock]), # only if still HELD
("STOCK_HELD", "PaymentDeclined"): ("REJECTED", [ReleaseStock, NotifyCustomer]),
("STOCK_HELD", "StockReleased"): ("EXPIRED", [VoidPaymentIfAny]),
("STOCK_HELD", "CancelRequested"): ("CANCELLED", [ReleaseStock, VoidPaymentIfAny]),
("EXPIRED", "PaymentAuthorized"): ("EXPIRED", [VoidPayment]), # lost the race
("CANCELLED", "PaymentAuthorized"): ("CANCELLED", [VoidPayment]),
("CONFIRMED", "StockCommitted"): ("CONFIRMED", [RequestFulfilment]),
("CONFIRMED", "StockCommitFailed"): ("REJECTED", [VoidPayment, NotifyCustomer]),
("CONFIRMED", "Shipped"): ("SHIPPED", [CapturePayment]),
("CONFIRMED", "CancelRequested"): ("CANCELLED", [ReleaseStock, VoidPayment]),
}The handler does three things atomically in the order database: checks the event was not already processed, applies a transition only if the table allows it from the current state, and queues the next commands in the outbox. An event that arrives in a state where it has no transition is logged and ignored rather than corrupting the order. That single rule absorbs most out-of-order and duplicate delivery.
Inventory: reservations with a time limit
Stock is where oversell happens. The inventory service keeps two counters per SKU, available and reserved, and a reservation row per order line. A reservation is a hold with an expiry, not a sale.
-- Reserve stock for one order line. Runs inside the inventory service's own database.
-- The unique key makes a redelivered ReserveStock a no-op instead of a second hold.
INSERT INTO reservations (order_id, sku, qty, expires_at, state)
SELECT :order_id, :sku, :qty, now() + interval '15 minutes', 'HELD'
WHERE NOT EXISTS (SELECT 1 FROM reservations WHERE order_id = :order_id AND sku = :sku);
UPDATE stock
SET available = available - :qty,
reserved = reserved + :qty
WHERE sku = :sku
AND available >= :qty; -- conditional update: never goes negative
-- 0 rows updated means insufficient stock: roll back and emit StockRejected.
-- Expiry sweeper, every minute: release holds whose order never paid.
UPDATE reservations SET state = 'EXPIRED'
WHERE state = 'HELD' AND expires_at < now()
RETURNING order_id, sku, qty; -- for each row: stock.available += qty, emit StockReleasedThe conditional update is what prevents oversell: two concurrent reservations for the last unit both run it, and the database guarantees only one sees a row updated. The unique reservation per order and SKU makes a redelivered command harmless. The expiry is what prevents stock being locked forever by abandoned checkouts or a payment service outage: after 15 minutes the sweeper releases the hold and emits StockReleased, and the order service moves the order to EXPIRED.
Pick the TTL from the payment path: it must exceed the longest realistic authorisation, including 3-D Secure challenges, plus retry budget. Too short releases stock under a paying customer; too long lets abandoned carts hold a flash sale's stock. A multi-line order reserves all lines or none: if line three is rejected, the inventory service releases lines one and two in the same handler before replying.
Payment: authorise, then capture
Card payments have two steps for good reason. Authorisation places a hold on the customer's funds; capture actually moves the money. Authorise at checkout and capture at shipment, so a cancelled or unshippable order needs a void rather than a refund. Check how long your payment provider keeps an authorisation valid; if fulfilment can take longer, you need re-authorisation logic.
Every call to the payment provider carries an idempotency key derived from the order and the operation, for example ord_8KQ2V7:authorize:1. When the call times out, the outcome is unknown, not failed. The payment service must not guess: it records the attempt as UNKNOWN, queries the provider by that key or reference, and replies only once it knows. Treating a timeout as a decline and retrying with a new key is how customers get charged twice.
Consumers, ordering and the read model
Every consumer, not only the order service, follows the same discipline: dedupe on event_id in the same transaction as the side effect, apply only transitions valid from the current state, and discard events whose order_version is not newer than what it holds. Delivery is at least once, so duplicates are normal operation.
Customers and support staff read order status from a separate read model built by a projection consumer, the pattern described in CQRS architecture. The projection lags, typically by tens to hundreds of milliseconds, and the confirmation page is where that shows. Serve it from the order service's own response, or poll the read model with the returned version until it catches up. If you later want a full audit history, the order events are close to an event-sourced log already; event sourcing covers what it would take to make them the source of truth.
Failure walk-throughs
| Failure | What happens | Why the design survives it |
|---|---|---|
| Relay crashes after publishing, before marking rows sent | the same events publish again | consumers dedupe on event_id |
| Inventory service down for 10 minutes | ReserveStock waits in the topic; orders stay PENDING | no stock was held; orders resume in order when it returns; the reconciler flags any past the SLO |
| Payment call times out | attempt recorded as UNKNOWN | status query by idempotency key decides; no blind retry |
| Reservation expiry races authorisation | StockReleased and PaymentAuthorized arrive close together | if EXPIRED wins, the late authorisation triggers VoidPayment; if CONFIRMED wins, CommitStock succeeds only on a still-HELD reservation, otherwise StockCommitFailed rejects the order and voids payment |
| Poison message a consumer cannot parse | handler fails repeatedly | after N attempts the message moves to a dead-letter topic with an alert; the partition keeps flowing |
| Rebalance mid-batch | a consumer reprocesses events it had applied | processed_events check makes the replay a no-op |
| Customer cancels while payment is in flight | CancelRequested arrives in STOCK_HELD | order is CANCELLED and stock released; a late PaymentAuthorized triggers VoidPayment |
Dead-lettering keeps other orders flowing but leaves that one order stuck, which is acceptable only because the reconciler below guarantees someone sees it.
Operating it
Watch four signals: outbox lag (age of the oldest unpublished row), consumer lag per partition (one hot partition usually means a poison message), dead-letter depth, which should page when non-zero, and orders by state and age. Orders in PENDING or STOCK_HELD for more than, say, 20 minutes is the most valuable dashboard in the system.
Back it with a reconciler: a scheduled job that asks inventory and payment for the facts about each stuck order and emits the missing event or the compensation. Events are the fast path; the reconciler makes the no-silent-stuck-order requirement true. Keep enough retention to replay a day into a rebuilt read model, and rehearse it.
Trade-offs
- Availability over immediacy. The API returns 202 and confirms asynchronously: resilient to downstream outages, but the UI must show pending states.
- Orchestrator coupling. The order service knows every checkout step, which suits a flow with money in it and costs teams who want to add steps independently.
- Operational surface. Brokers, relays, dead-letter topics and a reconciler are real systems to run. For a few orders per minute, one database with a job queue is simpler.
What to do next
- Write your order state table with every legal transition, and make the order service reject everything else.
- Define the event envelope with event_id, order_id, order_version, correlation and causation ids, and add a contract test to CI.
- Remove any broker call made inside a request transaction; route it through an outbox.
- Add processed-event dedupe to every consumer, in the same transaction as its side effect.
- Set the reservation TTL from measured payment-authorisation latency, and test the expiry-versus-authorisation race explicitly.
- Build the stuck-orders dashboard and the reconciler before launch, then rehearse a dead-letter replay and a read-model rebuild.