A payment system looks like a thin wrapper around a payment provider's API: send an amount, get back approved or declined. The hard part is everything that wrapper hides. The network can drop the answer after money has moved. Webhooks arrive twice, late, or out of order. Sellers must be paid days later, net of fees and refunds. Finance must be able to prove, every morning, that every cent the bank moved matches a record you hold.

This article designs a complete payment system for a two-sided marketplace: buyers pay the platform, the platform keeps a fee and pays sellers. It works through requirements and capacity, the data model, the payment state machine, how to call a provider safely, the ledger postings for a worked order, payouts, reconciliation, scaling and failure modes. Gateway internals such as card tokenization are covered in payment gateway architecture; here the focus is the system a product company builds around one or more gateways.

Advertisement

Requirements and a capacity estimate

Functional requirements: accept a payment for an order through at least one payment service provider (PSP); capture it when the order ships; refund fully or partially; handle disputes; pay each seller their balance on a schedule; and give finance daily reports that tie out to the bank.

Non-functional requirements matter more than usual. Correctness beats availability: a declined checkout is a lost sale, but a double charge is a support ticket, a chargeback and possibly a regulator. Every state change must be auditable months later. Retries must be safe end to end.

Capacity, worked through: 5 million orders a day is about 58 payments per second on average. Payment traffic is peaky, so plan for 10x, around 600 per second, at a sale. Each payment produces roughly one intent row, one or two attempt rows, three to five status events and two or three ledger transactions of two to four entries each, so around 15 writes per payment and 9,000 writes per second at peak. A relational cluster sharded a few ways handles that. Storage is modest: at roughly 2 KB per payment across all tables, 5 million payments a day is 10 GB a day, about 3.6 TB a year before indexes, and ledger data is kept for years for audit.

The architecture

The design splits into a small synchronous path and a larger asynchronous one. The checkout API calls the payment service, which records an intent, calls the provider through an adapter and returns. Everything else, including ledger postings, balances, payouts and reconciliation, is driven by events that the payment service writes to an outbox in the same database transaction as the state change, using the transactional outbox.

Marketplace payment system: synchronous path (top) and asynchronous money movement (bottom)Checkout / APIIdempotency-KeyPayment serviceintents + attemptsPSP adapterone per providerExternal PSPcard / wallet / bankauthorizewebhooksWebhook receiverdedupe by event idPayments DBstate + outboxLedger servicedouble-entry, appendBalancesderived per accountPayout schedulerholds, batches, railsReconciliationledger vs PSP vs bankoutbox eventsOnly the top row is on the customer's latency path; everything below is driven by events and daily files
The synchronous path records intent and attempt, calls the PSP and returns. State changes commit together with outbox events, which drive the ledger, balances and payouts; reconciliation compares the ledger with PSP and bank files.

Why separate a payment service from a ledger service? They change for different reasons. The payment service knows about providers, retries and states; the ledger knows only about accounts and balanced transactions. Keeping the ledger ignorant of providers means a new PSP never touches accounting code, and keeping it append-only means history can always be replayed. Double-entry ledger architecture goes deeper into the ledger itself.

One adapter per provider maps a common interface (authorize, capture, refund, query, parse webhook) to that provider's API, which makes routing and failover between providers possible.

Advertisement

The data model

Three tables carry the design. A payment intent is the customer's wish to pay for one order; it has exactly one current status. A payment attempt is one call to one provider; an intent can have several, for example a decline followed by a retry with another card. Ledger entries record money movement and are never updated.

-- amounts are integer minor units (cents, paise); never floats
CREATE TABLE payment_intent (
  id              uuid PRIMARY KEY,
  merchant_id     bigint      NOT NULL,           -- shard key
  order_id        text        NOT NULL,
  amount_minor    bigint      NOT NULL CHECK (amount_minor > 0),
  currency        char(3)     NOT NULL,
  status          text        NOT NULL,           -- see state machine
  idempotency_key text        NOT NULL,
  version         int         NOT NULL DEFAULT 0, -- optimistic concurrency
  created_at      timestamptz NOT NULL DEFAULT now(),
  UNIQUE (merchant_id, idempotency_key)
);

CREATE TABLE payment_attempt (                    -- one row per call to a PSP
  id           uuid PRIMARY KEY,
  intent_id    uuid NOT NULL REFERENCES payment_intent(id),
  psp          text NOT NULL,
  psp_ref      text,                              -- PSP's id, unknown until it answers
  status       text NOT NULL,                     -- SENT, SUCCEEDED, FAILED, UNKNOWN
  UNIQUE (psp, psp_ref)
);

CREATE TABLE ledger_txn (                         -- append-only: no UPDATE, no DELETE
  txn_id       uuid PRIMARY KEY,
  source_event text NOT NULL UNIQUE,              -- one txn per event: posting is idempotent
  posted_at    timestamptz NOT NULL DEFAULT now()
);

CREATE TABLE ledger_entry (                       -- entries of one txn sum to zero
  txn_id       uuid   NOT NULL REFERENCES ledger_txn(txn_id),
  account      text   NOT NULL,                   -- e.g. 'seller:42:payable'
  amount_minor bigint NOT NULL,                   -- debit positive, credit negative
  currency     char(3) NOT NULL,
  PRIMARY KEY (txn_id, account)
);

Three details are deliberate. Amounts are integers in minor units, because floating point cannot represent 0.10 exactly and rounding errors accumulate across millions of rows. The unique constraint on merchant and idempotency key makes a retried API request return the original intent instead of creating a second one; idempotency architecture covers key scope and expiry. Each ledger transaction header carries a unique source_event, and its entries are inserted in the same database transaction, so an event consumed twice cannot post money twice.

A payment state machine that cannot be corrupted

Status is not a free-form field that any code path may overwrite. It is a state machine with an explicit list of legal transitions, and every transition is a compare-and-set against the expected current state. Two workers racing on the same intent, for example a webhook handler and a timeout resolver, cannot both win: the second update matches zero rows and must re-read.

ALLOWED = {
    "REQUIRES_PAYMENT": {"PROCESSING"},
    "PROCESSING":       {"AUTHORIZED", "FAILED", "UNKNOWN"},
    "UNKNOWN":          {"AUTHORIZED", "FAILED"},        # resolved by query or webhook
    "AUTHORIZED":       {"CAPTURED", "VOIDED"},
    "CAPTURED":         {"PARTIALLY_REFUNDED", "REFUNDED", "DISPUTED"},
    "PARTIALLY_REFUNDED": {"PARTIALLY_REFUNDED", "REFUNDED", "DISPUTED"},
    "DISPUTED":         {"DISPUTE_WON", "DISPUTE_LOST"},
}

def transition(db, intent_id, expected, new, event):
    if new not in ALLOWED.get(expected, set()):
        raise IllegalTransition(expected, new)
    with db.transaction():
        rows = db.execute(
            "UPDATE payment_intent SET status=%s, version=version+1 "
            "WHERE id=%s AND status=%s", (new, intent_id, expected)).rowcount
        if rows == 0:                    # someone else moved it first: re-read, don't overwrite
            raise ConcurrentTransition(intent_id)
        db.execute("INSERT INTO outbox(topic, key, payload) VALUES (%s,%s,%s)",
                   ("payment.status", intent_id, event.to_json()))

The UNKNOWN state is the important one. It says honestly that the system sent a request and does not know the outcome. Collapsing it into FAILED is the classic bug: the customer retries, the first authorization turns out to have succeeded, and they are charged twice. Terminal states have no outgoing transitions, so late webhooks for them are logged and ignored.

Calling the provider safely

The rule is: persist before you call, and query before you retry. Write the attempt row first, so a crash mid-call leaves evidence. Pass the attempt's id to the provider as its idempotency key, where the provider supports one, so a retry of the same attempt cannot create a second charge. Use a timeout shorter than the checkout's own.

def authorize(intent):
    attempt = new_attempt(intent)                     # persisted BEFORE the network call
    try:
        resp = psp.authorize(amount=intent.amount_minor, currency=intent.currency,
                             idempotency_key=str(attempt.id), timeout=8.0)
    except (Timeout, ConnectionReset):
        mark_attempt(attempt, "UNKNOWN")
        transition(db, intent.id, "PROCESSING", "UNKNOWN", Event("psp_timeout"))
        schedule(resolve_unknown, attempt.id, delay=30)   # query-before-retry
        return "pending"
    if resp.approved:
        mark_attempt(attempt, "SUCCEEDED", psp_ref=resp.id)
        transition(db, intent.id, "PROCESSING", "AUTHORIZED", Event("authorized", resp.id))
    else:
        mark_attempt(attempt, "FAILED", psp_ref=resp.id)
        transition(db, intent.id, "PROCESSING", "FAILED", Event("declined", resp.decline_code))
    return resp.approved

When the call times out, resolve_unknown asks the provider for the attempt's status, using the idempotency key or your reference, with backoff. If the provider says it succeeded, transition to AUTHORIZED; if it never saw the request, it is safe to mark the attempt failed and let the customer try again. The provider's webhook may arrive first; that is fine, because both paths go through the same guarded transition.

Webhooks need the same discipline. Verify the signature, store the raw event keyed by the provider's event id before processing (duplicates then become a no-op), return 2xx quickly, and process asynchronously. Never assume order: a charge.refunded event can arrive before charge.captured is processed, so the handler must check the state machine and park events whose preconditions are not met yet.

Worked example: money movement for one order

A buyer pays 100.00 for an order from seller 42. The platform fee is 10%. The PSP charges 2.9% plus 0.30 per transaction and settles to the platform's bank account two days later. Seller 42 is paid weekly. Each ledger transaction below balances to zero; debits are positive.

EventAccountDebitCredit
Capture succeededpsp_receivable100.00
seller:42:payable90.00
platform:fee_revenue10.00
PSP settlement filebank:operating96.80
expense:psp_fees3.20
psp_receivable100.00
Weekly payoutseller:42:payable90.00
bank:operating90.00

Read the balances off the entries. After settlement, psp_receivable is back to zero, which is exactly what reconciliation checks: a non-zero receivable older than the settlement window means money the PSP owes you and has not paid. The platform earned 10.00, paid 3.20 in fees and keeps 6.80; the seller's payable is a liability until payout. Balances are sums of entries, cached in a projection that can be rebuilt at any time.

Now a refund of 30.00 after the payout. The buyer gets 30.00 back through the PSP, the platform reverses its fee proportionally (3.00), and seller 42 owes 27.00 they have already been paid. Their payable goes negative, and the next payout nets it off. Policy must cover the case where it never does, typically with a reserve held back from new sellers.

Payouts, holds and multi-step flows

Pay-in and pay-out are different problems. Pay-in is many small, latency-sensitive, customer-initiated calls. Pay-out is few, large, scheduled bank transfers where the risk is paying the wrong amount or paying twice. A payout run takes a consistent snapshot of payable balances, subtracts holds (disputed funds, a rolling reserve for new sellers, anything flagged by risk), creates one payout record per seller with a deterministic idempotency key such as seller_id + payout_date, posts the ledger transaction, and only then submits to the bank rail. If the bank rejects the transfer, a reversing transaction restores the payable.

An order that spans several services, for example reserve inventory, authorize, then capture on shipment, is a long-running workflow with compensations rather than a distributed transaction: if inventory reservation fails after authorization, the compensation is a void. Saga pattern architecture explains orchestration versus choreography for these flows; for payments, an orchestrator with an explicit, persisted workflow state is easier to audit.

Reconciliation as a daily process

Reconciliation is three-way. Your ledger says what should have happened. The PSP's settlement report says what the provider did and what it paid you. The bank statement says what actually arrived. A daily job matches records across all three by provider reference and amount, and classifies every mismatch as a break.

  • In ledger, missing at PSP: usually an UNKNOWN attempt that actually failed, or a capture that was never sent.
  • At PSP, missing in ledger: the dangerous one. Money moved without a record, typically a lost webhook or a bug that skipped posting.
  • Amount mismatch: fee differences, currency conversion, or partial captures.
  • Timing: settled after the cut-off; expected to clear the next day, and an alert if it does not.

Measure breaks by count and value per day and age them. A healthy system has a small, stable number of timing breaks and near zero of the others. Every break must end in a correcting ledger transaction with a reason, never an edited row.

Scaling, failure modes and trade-offs

Shard payment data by merchant, so one merchant's intents, attempts and payables live together and transitions stay single-shard. Ledger transactions that touch a merchant account and a platform account cross shards; either keep platform accounts on every shard as sub-accounts that roll up, or post platform entries asynchronously from the outbox and let reconciliation prove they balance. The first keeps every transaction atomic, which is usually worth the extra roll-up job.

FailureWhat goes wrongDefence
PSP timeoutMoney moved, record says failed; retry double chargesUNKNOWN state, query-before-retry, provider idempotency key
Duplicate webhook or eventDouble posting to ledgerUnique event id on receipt and on the ledger transaction's source_event
Crash between DB write and publishLedger never learns of captureTransactional outbox, not dual writes
Concurrent handlersStatus overwritten backwardsCompare-and-set transitions
Payout job rerunSeller paid twiceDeterministic payout idempotency key, ledger posted before bank submit
PSP outageCheckout failsSecond provider behind the adapter; route by health

The main trade-off is synchronous certainty versus availability. Holding checkout open until the provider answers gives the customer a definite result but ties your latency to the provider's. Returning pending and confirming later keeps checkout fast but needs good UX and order-hold logic. Most marketplaces authorize synchronously and do everything else asynchronously.

What to do next

  1. Write the state machine for your payments as a table of legal transitions, including an explicit UNKNOWN state, and enforce it with compare-and-set updates.
  2. Store money as integer minor units with a currency column, and add unique constraints for API idempotency keys, provider references and ledger source events.
  3. Persist every attempt before calling the provider, pass a provider idempotency key, and implement query-before-retry for timeouts.
  4. Post the worked example above through your ledger in a test and assert that every transaction sums to zero and the receivable returns to zero after settlement.
  5. Build the daily three-way reconciliation job before launch, with break categories, values and ageing on a dashboard.
  6. Run a game day: kill the service mid-authorization, replay a webhook batch twice, and rerun a payout job; confirm no money is duplicated or lost.
Key takeaway: A payment system is a state machine plus a ledger, glued together by idempotency. Record intent and attempt before calling a provider, treat timeouts as UNKNOWN and resolve them by querying, and drive every state change through guarded transitions that also write outbox events. Post money movement as balanced, append-only ledger transactions, pay sellers from derived balances with deterministic payout keys, and prove it all daily with three-way reconciliation. Scale by sharding on merchant, and keep the synchronous path limited to what the customer must wait for.