A payment system looks like a thin wrapper around a payment provider's API: send an amount, get back approved or declined. The hard part is everything that wrapper hides. The network can drop the answer after money has moved. Webhooks arrive twice, late, or out of order. Sellers must be paid days later, net of fees and refunds. Finance must be able to prove, every morning, that every cent the bank moved matches a record you hold.
This article designs a complete payment system for a two-sided marketplace: buyers pay the platform, the platform keeps a fee and pays sellers. It works through requirements and capacity, the data model, the payment state machine, how to call a provider safely, the ledger postings for a worked order, payouts, reconciliation, scaling and failure modes. Gateway internals such as card tokenization are covered in payment gateway architecture; here the focus is the system a product company builds around one or more gateways.
Requirements and a capacity estimate
Functional requirements: accept a payment for an order through at least one payment service provider (PSP); capture it when the order ships; refund fully or partially; handle disputes; pay each seller their balance on a schedule; and give finance daily reports that tie out to the bank.
Non-functional requirements matter more than usual. Correctness beats availability: a declined checkout is a lost sale, but a double charge is a support ticket, a chargeback and possibly a regulator. Every state change must be auditable months later. Retries must be safe end to end.
Capacity, worked through: 5 million orders a day is about 58 payments per second on average. Payment traffic is peaky, so plan for 10x, around 600 per second, at a sale. Each payment produces roughly one intent row, one or two attempt rows, three to five status events and two or three ledger transactions of two to four entries each, so around 15 writes per payment and 9,000 writes per second at peak. A relational cluster sharded a few ways handles that. Storage is modest: at roughly 2 KB per payment across all tables, 5 million payments a day is 10 GB a day, about 3.6 TB a year before indexes, and ledger data is kept for years for audit.
The architecture
The design splits into a small synchronous path and a larger asynchronous one. The checkout API calls the payment service, which records an intent, calls the provider through an adapter and returns. Everything else, including ledger postings, balances, payouts and reconciliation, is driven by events that the payment service writes to an outbox in the same database transaction as the state change, using the transactional outbox.
Why separate a payment service from a ledger service? They change for different reasons. The payment service knows about providers, retries and states; the ledger knows only about accounts and balanced transactions. Keeping the ledger ignorant of providers means a new PSP never touches accounting code, and keeping it append-only means history can always be replayed. Double-entry ledger architecture goes deeper into the ledger itself.
One adapter per provider maps a common interface (authorize, capture, refund, query, parse webhook) to that provider's API, which makes routing and failover between providers possible.
The data model
Three tables carry the design. A payment intent is the customer's wish to pay for one order; it has exactly one current status. A payment attempt is one call to one provider; an intent can have several, for example a decline followed by a retry with another card. Ledger entries record money movement and are never updated.
-- amounts are integer minor units (cents, paise); never floats
CREATE TABLE payment_intent (
id uuid PRIMARY KEY,
merchant_id bigint NOT NULL, -- shard key
order_id text NOT NULL,
amount_minor bigint NOT NULL CHECK (amount_minor > 0),
currency char(3) NOT NULL,
status text NOT NULL, -- see state machine
idempotency_key text NOT NULL,
version int NOT NULL DEFAULT 0, -- optimistic concurrency
created_at timestamptz NOT NULL DEFAULT now(),
UNIQUE (merchant_id, idempotency_key)
);
CREATE TABLE payment_attempt ( -- one row per call to a PSP
id uuid PRIMARY KEY,
intent_id uuid NOT NULL REFERENCES payment_intent(id),
psp text NOT NULL,
psp_ref text, -- PSP's id, unknown until it answers
status text NOT NULL, -- SENT, SUCCEEDED, FAILED, UNKNOWN
UNIQUE (psp, psp_ref)
);
CREATE TABLE ledger_txn ( -- append-only: no UPDATE, no DELETE
txn_id uuid PRIMARY KEY,
source_event text NOT NULL UNIQUE, -- one txn per event: posting is idempotent
posted_at timestamptz NOT NULL DEFAULT now()
);
CREATE TABLE ledger_entry ( -- entries of one txn sum to zero
txn_id uuid NOT NULL REFERENCES ledger_txn(txn_id),
account text NOT NULL, -- e.g. 'seller:42:payable'
amount_minor bigint NOT NULL, -- debit positive, credit negative
currency char(3) NOT NULL,
PRIMARY KEY (txn_id, account)
);Three details are deliberate. Amounts are integers in minor units, because floating point cannot represent 0.10 exactly and rounding errors accumulate across millions of rows. The unique constraint on merchant and idempotency key makes a retried API request return the original intent instead of creating a second one; idempotency architecture covers key scope and expiry. Each ledger transaction header carries a unique source_event, and its entries are inserted in the same database transaction, so an event consumed twice cannot post money twice.
A payment state machine that cannot be corrupted
Status is not a free-form field that any code path may overwrite. It is a state machine with an explicit list of legal transitions, and every transition is a compare-and-set against the expected current state. Two workers racing on the same intent, for example a webhook handler and a timeout resolver, cannot both win: the second update matches zero rows and must re-read.
ALLOWED = {
"REQUIRES_PAYMENT": {"PROCESSING"},
"PROCESSING": {"AUTHORIZED", "FAILED", "UNKNOWN"},
"UNKNOWN": {"AUTHORIZED", "FAILED"}, # resolved by query or webhook
"AUTHORIZED": {"CAPTURED", "VOIDED"},
"CAPTURED": {"PARTIALLY_REFUNDED", "REFUNDED", "DISPUTED"},
"PARTIALLY_REFUNDED": {"PARTIALLY_REFUNDED", "REFUNDED", "DISPUTED"},
"DISPUTED": {"DISPUTE_WON", "DISPUTE_LOST"},
}
def transition(db, intent_id, expected, new, event):
if new not in ALLOWED.get(expected, set()):
raise IllegalTransition(expected, new)
with db.transaction():
rows = db.execute(
"UPDATE payment_intent SET status=%s, version=version+1 "
"WHERE id=%s AND status=%s", (new, intent_id, expected)).rowcount
if rows == 0: # someone else moved it first: re-read, don't overwrite
raise ConcurrentTransition(intent_id)
db.execute("INSERT INTO outbox(topic, key, payload) VALUES (%s,%s,%s)",
("payment.status", intent_id, event.to_json()))The UNKNOWN state is the important one. It says honestly that the system sent a request and does not know the outcome. Collapsing it into FAILED is the classic bug: the customer retries, the first authorization turns out to have succeeded, and they are charged twice. Terminal states have no outgoing transitions, so late webhooks for them are logged and ignored.
Calling the provider safely
The rule is: persist before you call, and query before you retry. Write the attempt row first, so a crash mid-call leaves evidence. Pass the attempt's id to the provider as its idempotency key, where the provider supports one, so a retry of the same attempt cannot create a second charge. Use a timeout shorter than the checkout's own.
def authorize(intent):
attempt = new_attempt(intent) # persisted BEFORE the network call
try:
resp = psp.authorize(amount=intent.amount_minor, currency=intent.currency,
idempotency_key=str(attempt.id), timeout=8.0)
except (Timeout, ConnectionReset):
mark_attempt(attempt, "UNKNOWN")
transition(db, intent.id, "PROCESSING", "UNKNOWN", Event("psp_timeout"))
schedule(resolve_unknown, attempt.id, delay=30) # query-before-retry
return "pending"
if resp.approved:
mark_attempt(attempt, "SUCCEEDED", psp_ref=resp.id)
transition(db, intent.id, "PROCESSING", "AUTHORIZED", Event("authorized", resp.id))
else:
mark_attempt(attempt, "FAILED", psp_ref=resp.id)
transition(db, intent.id, "PROCESSING", "FAILED", Event("declined", resp.decline_code))
return resp.approvedWhen the call times out, resolve_unknown asks the provider for the attempt's status, using the idempotency key or your reference, with backoff. If the provider says it succeeded, transition to AUTHORIZED; if it never saw the request, it is safe to mark the attempt failed and let the customer try again. The provider's webhook may arrive first; that is fine, because both paths go through the same guarded transition.
Webhooks need the same discipline. Verify the signature, store the raw event keyed by the provider's event id before processing (duplicates then become a no-op), return 2xx quickly, and process asynchronously. Never assume order: a charge.refunded event can arrive before charge.captured is processed, so the handler must check the state machine and park events whose preconditions are not met yet.
Worked example: money movement for one order
A buyer pays 100.00 for an order from seller 42. The platform fee is 10%. The PSP charges 2.9% plus 0.30 per transaction and settles to the platform's bank account two days later. Seller 42 is paid weekly. Each ledger transaction below balances to zero; debits are positive.
| Event | Account | Debit | Credit |
|---|---|---|---|
| Capture succeeded | psp_receivable | 100.00 | |
| seller:42:payable | 90.00 | ||
| platform:fee_revenue | 10.00 | ||
| PSP settlement file | bank:operating | 96.80 | |
| expense:psp_fees | 3.20 | ||
| psp_receivable | 100.00 | ||
| Weekly payout | seller:42:payable | 90.00 | |
| bank:operating | 90.00 |
Read the balances off the entries. After settlement, psp_receivable is back to zero, which is exactly what reconciliation checks: a non-zero receivable older than the settlement window means money the PSP owes you and has not paid. The platform earned 10.00, paid 3.20 in fees and keeps 6.80; the seller's payable is a liability until payout. Balances are sums of entries, cached in a projection that can be rebuilt at any time.
Now a refund of 30.00 after the payout. The buyer gets 30.00 back through the PSP, the platform reverses its fee proportionally (3.00), and seller 42 owes 27.00 they have already been paid. Their payable goes negative, and the next payout nets it off. Policy must cover the case where it never does, typically with a reserve held back from new sellers.
Payouts, holds and multi-step flows
Pay-in and pay-out are different problems. Pay-in is many small, latency-sensitive, customer-initiated calls. Pay-out is few, large, scheduled bank transfers where the risk is paying the wrong amount or paying twice. A payout run takes a consistent snapshot of payable balances, subtracts holds (disputed funds, a rolling reserve for new sellers, anything flagged by risk), creates one payout record per seller with a deterministic idempotency key such as seller_id + payout_date, posts the ledger transaction, and only then submits to the bank rail. If the bank rejects the transfer, a reversing transaction restores the payable.
An order that spans several services, for example reserve inventory, authorize, then capture on shipment, is a long-running workflow with compensations rather than a distributed transaction: if inventory reservation fails after authorization, the compensation is a void. Saga pattern architecture explains orchestration versus choreography for these flows; for payments, an orchestrator with an explicit, persisted workflow state is easier to audit.
Reconciliation as a daily process
Reconciliation is three-way. Your ledger says what should have happened. The PSP's settlement report says what the provider did and what it paid you. The bank statement says what actually arrived. A daily job matches records across all three by provider reference and amount, and classifies every mismatch as a break.
- In ledger, missing at PSP: usually an
UNKNOWNattempt that actually failed, or a capture that was never sent. - At PSP, missing in ledger: the dangerous one. Money moved without a record, typically a lost webhook or a bug that skipped posting.
- Amount mismatch: fee differences, currency conversion, or partial captures.
- Timing: settled after the cut-off; expected to clear the next day, and an alert if it does not.
Measure breaks by count and value per day and age them. A healthy system has a small, stable number of timing breaks and near zero of the others. Every break must end in a correcting ledger transaction with a reason, never an edited row.
Scaling, failure modes and trade-offs
Shard payment data by merchant, so one merchant's intents, attempts and payables live together and transitions stay single-shard. Ledger transactions that touch a merchant account and a platform account cross shards; either keep platform accounts on every shard as sub-accounts that roll up, or post platform entries asynchronously from the outbox and let reconciliation prove they balance. The first keeps every transaction atomic, which is usually worth the extra roll-up job.
| Failure | What goes wrong | Defence |
|---|---|---|
| PSP timeout | Money moved, record says failed; retry double charges | UNKNOWN state, query-before-retry, provider idempotency key |
| Duplicate webhook or event | Double posting to ledger | Unique event id on receipt and on the ledger transaction's source_event |
| Crash between DB write and publish | Ledger never learns of capture | Transactional outbox, not dual writes |
| Concurrent handlers | Status overwritten backwards | Compare-and-set transitions |
| Payout job rerun | Seller paid twice | Deterministic payout idempotency key, ledger posted before bank submit |
| PSP outage | Checkout fails | Second provider behind the adapter; route by health |
The main trade-off is synchronous certainty versus availability. Holding checkout open until the provider answers gives the customer a definite result but ties your latency to the provider's. Returning pending and confirming later keeps checkout fast but needs good UX and order-hold logic. Most marketplaces authorize synchronously and do everything else asynchronously.
What to do next
- Write the state machine for your payments as a table of legal transitions, including an explicit UNKNOWN state, and enforce it with compare-and-set updates.
- Store money as integer minor units with a currency column, and add unique constraints for API idempotency keys, provider references and ledger source events.
- Persist every attempt before calling the provider, pass a provider idempotency key, and implement query-before-retry for timeouts.
- Post the worked example above through your ledger in a test and assert that every transaction sums to zero and the receivable returns to zero after settlement.
- Build the daily three-way reconciliation job before launch, with break categories, values and ageing on a dashboard.
- Run a game day: kill the service mid-authorization, replay a webhook batch twice, and rerun a payout job; confirm no money is duplicated or lost.