Most write-ups of event-driven architecture stop at the pattern names: outbox, saga, idempotency, CQRS. An order system is where you discover how they fit together, because it touches money, physical stock and a waiting customer. Get the fit wrong and you oversell the last unit, charge a card twice, or strand an order in a state nobody owns.

This article designs one order system end to end: the order as a state machine, the events that move it, the checkout path through inventory and payment with code, and the failures that actually happen. The underlying mechanisms have their own articles on this site, linked where they appear; here the focus is the domain design that uses them.

Advertisement

Scope and a worked load

The system takes orders for a mid-sized online retailer. Customers place an order containing one or more stock-keeping units (SKUs); the system must hold stock, authorise payment, hand the order to the warehouse, capture payment on shipment, and let the customer see status at every step. Cancellation is allowed until the order ships.

Assume 1,500 orders per minute normally, 20,000 in the first minute of a flash sale, and about 12 events per order across its life. The flash peak is roughly 333 orders and 4,000 events per second, which any modern broker handles. The hard problems are not throughput but correctness under retries, ordering, and partial failure across services that each own a database.

Three requirements drive everything: no oversell beyond an agreed tolerance, no double charge ever, and no silent stuck order: every order is either progressing or visible to someone as stuck.

The order is a state machine first

Before choosing topics or services, write the order's states and the only legal transitions between them. Most production bugs in order systems are illegal transitions nobody forbade.

StateEntered whenLegal next states
PENDINGorder accepted and persistedSTOCK_HELD, REJECTED, CANCELLED
STOCK_HELDall lines reservedCONFIRMED, REJECTED, EXPIRED, CANCELLED
CONFIRMEDpayment authorisedSHIPPED, CANCELLED, REJECTED
SHIPPEDcarrier accepted the parcelDELIVERED, RETURNED
REJECTED / EXPIRED / CANCELLEDstock, payment or customer said noterminal

Two choices follow. First, only the order service decides the order's state; other services own their facts (stock, payments, shipments) and report them as events. Second, each order carries a monotonically increasing version. Every state change increments it and every event the order service emits carries it, so consumers can tell a stale event from a fresh one without trusting delivery order.

Advertisement

Event contracts

Every event shares an envelope. The business payload sits in data; everything a consumer needs for deduplication, ordering and tracing sits outside it.

{
  "event_id":       "0192f1c4-6a0e-7c11-9d3a-5b2f8e41a7d0",
  "event_type":     "OrderPlaced",
  "schema_version": 3,
  "order_id":       "ord_8KQ2V7",
  "order_version":  1,
  "occurred_at":    "2026-09-30T08:05:12.418Z",
  "correlation_id": "chk_51ZP",
  "causation_id":   "cmd_PlaceOrder_51ZP",
  "producer":       "order-service",
  "data": {
    "customer_id": "cus_1934",
    "currency": "EUR",
    "lines": [{"sku": "SKU-4411", "qty": 2, "unit_price_minor": 1999}],
    "total_minor": 3998
  }
}

event_id is globally unique and is the deduplication key. order_id is the partition key. order_version lets a consumer discard an event older than what it has already applied. correlation_id ties every event of one checkout together for tracing; causation_id names the command or event that caused this one, which is what you need when reconstructing why an order did something odd. Money is carried in integer minor units with an explicit currency, never as floating point.

Carry the facts a consumer needs to act without calling back: fulfilment needs lines and address, notifications need customer and total. Callbacks reintroduce coupling and read the current state rather than the state at the time of the event. Evolve schemas additively, bump schema_version on any change, and keep a contract test in CI.

Architecture and the checkout data flow

Checkout path: one local transaction, then events keyed by order_idClientIdempotency-KeyOrder servicestate machineorders tableoutbox tablesame DB transactionOutbox relayCDC or pollerPOSTBroker: order-eventspartition = hash(order_id)Inventoryreserve with TTLPaymentauthorise, captureFulfilmentpick, pack, shipRead modelstatus viewsNotifyemail, pushDashed: replies (StockReserved, PaymentAuthorized, ...) flow back to the order service, which advances the stateEvery consumer is idempotent on event_id; ordering is only guaranteed within one order_id
The checkout path. The order row and its outgoing events commit in one local transaction; a relay publishes them to a topic partitioned by order_id; services act and reply; the order service advances the state machine.
  1. The client sends POST /orders with an Idempotency-Key header. A retried request with the same key returns the original response instead of creating a second order; the mechanics are in idempotency architecture.
  2. The order service inserts the order in PENDING and, in the same database transaction, inserts OrderPlaced and a ReserveStock command into an outbox table. It returns 202 Accepted with the order id. No broker call happens inside the request; that dual write is the classic bug the transactional outbox exists to remove.
  3. A relay, either change data capture on the outbox table or a poller, publishes the rows to the broker with order_id as the message key.
  4. Inventory reserves stock and replies StockReserved or StockRejected. The order service moves to STOCK_HELD and emits AuthorizePayment.
  5. Payment authorises the card and replies. On PaymentAuthorized the order becomes CONFIRMED and emits CommitStock, which turns the hold into a sale only if it is still HELD; fulfilment is requested on StockCommitted. Capture happens later, on Shipped.

Because the key is the order id, every event for one order lands on one partition and is consumed in the order it was published. That is the only ordering guarantee you get and the only one you need: nothing in this design depends on ordering across orders.

Orchestrate checkout, choreograph the rest

A saga can be choreographed, with each service reacting to the previous one's event, or orchestrated, with one component issuing commands and reacting to replies. The trade-offs in general are in saga pattern architecture. For orders the answer is usually both, split by purpose.

The checkout path is orchestrated by the order service, as the handler below shows. It has a strict sequence, compensations that depend on how far it got, and a timeout per step, and that logic belongs in one place you can read and test. Downstream reactions that do not affect the order's outcome, such as emails, loyalty points, analytics and search indexing, are choreographed: they subscribe to OrderConfirmed or OrderShipped and nobody waits for them.

# Order service: the checkout process manager. One handler per reply event.
def handle(event, db, outbox):
    with db.transaction():
        if db.exists("processed_events", event_id=event.event_id):
            return                                    # duplicate delivery: already applied
        order = db.select_for_update("orders", order_id=event.order_id)
        next_state, commands = TRANSITIONS.get((order.state, event.event_type), (None, []))
        if next_state is None:
            log.warning("ignored %s in state %s", event.event_type, order.state)
        else:
            order.state = next_state
            order.version += 1
            db.update("orders", order)
            for cmd in commands:                      # e.g. AuthorizePayment, ReleaseStock
                outbox.append(cmd.for_order(order))   # published later by the relay
        db.insert("processed_events", event_id=event.event_id)

TRANSITIONS = {
    ("PENDING",          "StockReserved"):     ("STOCK_HELD",  [AuthorizePayment]),
    ("PENDING",          "StockRejected"):     ("REJECTED",    [NotifyCustomer]),
    ("PENDING",          "CancelRequested"):   ("CANCELLED",   [ReleaseStock]),  # no-op if none held
    ("STOCK_HELD",       "PaymentAuthorized"): ("CONFIRMED",   [CommitStock]),   # only if still HELD
    ("STOCK_HELD",       "PaymentDeclined"):   ("REJECTED",    [ReleaseStock, NotifyCustomer]),
    ("STOCK_HELD",       "StockReleased"):     ("EXPIRED",     [VoidPaymentIfAny]),
    ("STOCK_HELD",       "CancelRequested"):   ("CANCELLED",   [ReleaseStock, VoidPaymentIfAny]),
    ("EXPIRED",          "PaymentAuthorized"): ("EXPIRED",     [VoidPayment]),   # lost the race
    ("CANCELLED",        "PaymentAuthorized"): ("CANCELLED",   [VoidPayment]),
    ("CONFIRMED",        "StockCommitted"):    ("CONFIRMED",   [RequestFulfilment]),
    ("CONFIRMED",        "StockCommitFailed"): ("REJECTED",    [VoidPayment, NotifyCustomer]),
    ("CONFIRMED",        "Shipped"):           ("SHIPPED",     [CapturePayment]),
    ("CONFIRMED",        "CancelRequested"):   ("CANCELLED",   [ReleaseStock, VoidPayment]),
}

The handler does three things atomically in the order database: checks the event was not already processed, applies a transition only if the table allows it from the current state, and queues the next commands in the outbox. An event that arrives in a state where it has no transition is logged and ignored rather than corrupting the order. That single rule absorbs most out-of-order and duplicate delivery.

Inventory: reservations with a time limit

Stock is where oversell happens. The inventory service keeps two counters per SKU, available and reserved, and a reservation row per order line. A reservation is a hold with an expiry, not a sale.

-- Reserve stock for one order line. Runs inside the inventory service's own database.
-- The unique key makes a redelivered ReserveStock a no-op instead of a second hold.
INSERT INTO reservations (order_id, sku, qty, expires_at, state)
SELECT :order_id, :sku, :qty, now() + interval '15 minutes', 'HELD'
WHERE NOT EXISTS (SELECT 1 FROM reservations WHERE order_id = :order_id AND sku = :sku);

UPDATE stock
   SET available = available - :qty,
       reserved  = reserved  + :qty
 WHERE sku = :sku
   AND available >= :qty;          -- conditional update: never goes negative
-- 0 rows updated means insufficient stock: roll back and emit StockRejected.

-- Expiry sweeper, every minute: release holds whose order never paid.
UPDATE reservations SET state = 'EXPIRED'
 WHERE state = 'HELD' AND expires_at < now()
RETURNING order_id, sku, qty;      -- for each row: stock.available += qty, emit StockReleased

The conditional update is what prevents oversell: two concurrent reservations for the last unit both run it, and the database guarantees only one sees a row updated. The unique reservation per order and SKU makes a redelivered command harmless. The expiry is what prevents stock being locked forever by abandoned checkouts or a payment service outage: after 15 minutes the sweeper releases the hold and emits StockReleased, and the order service moves the order to EXPIRED.

Pick the TTL from the payment path: it must exceed the longest realistic authorisation, including 3-D Secure challenges, plus retry budget. Too short releases stock under a paying customer; too long lets abandoned carts hold a flash sale's stock. A multi-line order reserves all lines or none: if line three is rejected, the inventory service releases lines one and two in the same handler before replying.

Payment: authorise, then capture

Card payments have two steps for good reason. Authorisation places a hold on the customer's funds; capture actually moves the money. Authorise at checkout and capture at shipment, so a cancelled or unshippable order needs a void rather than a refund. Check how long your payment provider keeps an authorisation valid; if fulfilment can take longer, you need re-authorisation logic.

Every call to the payment provider carries an idempotency key derived from the order and the operation, for example ord_8KQ2V7:authorize:1. When the call times out, the outcome is unknown, not failed. The payment service must not guess: it records the attempt as UNKNOWN, queries the provider by that key or reference, and replies only once it knows. Treating a timeout as a decline and retrying with a new key is how customers get charged twice.

Consumers, ordering and the read model

Every consumer, not only the order service, follows the same discipline: dedupe on event_id in the same transaction as the side effect, apply only transitions valid from the current state, and discard events whose order_version is not newer than what it holds. Delivery is at least once, so duplicates are normal operation.

Customers and support staff read order status from a separate read model built by a projection consumer, the pattern described in CQRS architecture. The projection lags, typically by tens to hundreds of milliseconds, and the confirmation page is where that shows. Serve it from the order service's own response, or poll the read model with the returned version until it catches up. If you later want a full audit history, the order events are close to an event-sourced log already; event sourcing covers what it would take to make them the source of truth.

Failure walk-throughs

FailureWhat happensWhy the design survives it
Relay crashes after publishing, before marking rows sentthe same events publish againconsumers dedupe on event_id
Inventory service down for 10 minutesReserveStock waits in the topic; orders stay PENDINGno stock was held; orders resume in order when it returns; the reconciler flags any past the SLO
Payment call times outattempt recorded as UNKNOWNstatus query by idempotency key decides; no blind retry
Reservation expiry races authorisationStockReleased and PaymentAuthorized arrive close togetherif EXPIRED wins, the late authorisation triggers VoidPayment; if CONFIRMED wins, CommitStock succeeds only on a still-HELD reservation, otherwise StockCommitFailed rejects the order and voids payment
Poison message a consumer cannot parsehandler fails repeatedlyafter N attempts the message moves to a dead-letter topic with an alert; the partition keeps flowing
Rebalance mid-batcha consumer reprocesses events it had appliedprocessed_events check makes the replay a no-op
Customer cancels while payment is in flightCancelRequested arrives in STOCK_HELDorder is CANCELLED and stock released; a late PaymentAuthorized triggers VoidPayment

Dead-lettering keeps other orders flowing but leaves that one order stuck, which is acceptable only because the reconciler below guarantees someone sees it.

Operating it

Watch four signals: outbox lag (age of the oldest unpublished row), consumer lag per partition (one hot partition usually means a poison message), dead-letter depth, which should page when non-zero, and orders by state and age. Orders in PENDING or STOCK_HELD for more than, say, 20 minutes is the most valuable dashboard in the system.

Back it with a reconciler: a scheduled job that asks inventory and payment for the facts about each stuck order and emits the missing event or the compensation. Events are the fast path; the reconciler makes the no-silent-stuck-order requirement true. Keep enough retention to replay a day into a rebuilt read model, and rehearse it.

Trade-offs

  • Availability over immediacy. The API returns 202 and confirms asynchronously: resilient to downstream outages, but the UI must show pending states.
  • Orchestrator coupling. The order service knows every checkout step, which suits a flow with money in it and costs teams who want to add steps independently.
  • Operational surface. Brokers, relays, dead-letter topics and a reconciler are real systems to run. For a few orders per minute, one database with a job queue is simpler.

What to do next

  1. Write your order state table with every legal transition, and make the order service reject everything else.
  2. Define the event envelope with event_id, order_id, order_version, correlation and causation ids, and add a contract test to CI.
  3. Remove any broker call made inside a request transaction; route it through an outbox.
  4. Add processed-event dedupe to every consumer, in the same transaction as its side effect.
  5. Set the reservation TTL from measured payment-authorisation latency, and test the expiry-versus-authorisation race explicitly.
  6. Build the stuck-orders dashboard and the reconciler before launch, then rehearse a dead-letter replay and a read-model rebuild.
Key takeaway: An event-driven order system is a state machine owned by one service, moved by versioned events keyed by order id. Commit state and outgoing events together through an outbox, reserve stock with a conditional update and a TTL, authorise at checkout and capture at shipment, and make every consumer idempotent on event id. Then assume events will be late, duplicated or lost to a dead-letter topic, and back the whole flow with a reconciler that finds and resolves stuck orders.