A shopping cart looks like the simplest feature in e-commerce: a list of products and quantities. At Amazon's scale it became one of the most influential designs in distributed systems, because it forced a question most systems avoid: when the network fails, is it better to refuse a customer's write or to accept it and sort out conflicts later? Amazon's answer, published in the 2007 Dynamo paper, was that an add-to-cart must essentially never fail.

This article builds a cart service from that requirement. It separates what is public (the Dynamo paper) from what is general engineering practice, because Amazon does not publish its current cart internals and nothing here should be read as a description of them. Dynamo is also not DynamoDB: the managed database shares a name and some ideas, not the replication design described below. You will leave with a data model, merge rules, a concurrency strategy, a checkout protocol and a list of failure modes to test.

Advertisement

Requirements that shape everything

Start with what the cart is for. It is a customer's intent to buy, not a reservation: adding an item does not hold stock or lock a price. That single decision removes most coordination from the hot path. The remaining requirements are:

  • Always writable. A failed add-to-cart is lost revenue and cannot be retried later by the business. Availability for writes beats consistency.
  • Tail latency. The Dynamo paper describes service-level targets at the 99.9th percentile, with 300 ms as its example, because averages hide the customers who give up.
  • Durable across devices. An item added on a phone must appear on a laptop, and must survive a server or data-centre failure.
  • Read-heavy composition. Each cart view shows current price, availability, delivery estimates and promotions, all owned by other services.
  • Correct at checkout. Whatever looseness the cart tolerates, the order must be built from a validated snapshot.

These pull in different directions. The design resolves them by being loose where looseness is cheap (the cart contents) and strict where it is expensive (checkout).

The architecture at a glance

Web / app clientcart token or sessionEdge + API layerauth, rate limitswriteCart servicestateless, mergesput / getReplicated cart storeN replicas, always writablepreference list: A, B, C (+ hinted handoff)renderCart page composerfan-out with deadlinesPricingcurrent priceAvailabilitystock, deliveryPromotionscoupons, dealsproceedCheckoutsnapshot + keyOrder systemreserves, chargesclear by versionCartdelete
A stateless cart service writes to a replicated store that stays writable during failures. The cart page is composed at read time from pricing, availability and promotions. Checkout takes a versioned snapshot; the cart is cleared only after the order system accepts the order.

The cart service is stateless and horizontally scaled. Its store holds only what the customer chose: SKU or offer identifier, quantity, when it was added and the price seen at that moment for display. Everything that changes independently of the customer, such as current price and stock, is fetched when the cart is rendered and again at checkout. Keeping those out of the stored cart means a price change never needs a write to millions of carts.

Advertisement

What the Dynamo paper tells us

The paper uses the cart as its motivating example. Keys are placed on a consistent-hash ring; each key is stored on N nodes, its preference list. A read waits for R replies and a write for W acknowledgements, and the paper cites (N, R, W) = (3, 2, 2) as a common configuration. Because R + W is greater than N, a read normally overlaps the latest write; quorum systems explains the arithmetic.

The important part is what happens during failures. Dynamo uses a sloppy quorum: if a preferred node is unreachable, the write goes to the next healthy node on the ring, which keeps it with a hint and hands it back when the owner recovers. That is hinted handoff, and it is why a cart write succeeds even when two of its three home replicas are down. The price is that two partitions can each accept writes to the same cart.

Dynamo tracks those divergent versions with vector clocks, as in vector clocks in depth. When one version descends from another, the older one is discarded. When they are concurrent, the store returns both, and the application reconciles them. For the cart, reconciliation merged the versions. The paper is candid about the consequence: an add is never lost, but a deleted item can reappear.

Worked example: why a naive merge resurrects deletions

Suppose a cart holds a charger and a case, version V1. A partition splits the replicas. On the phone the customer removes the case, producing V2 with only the charger. On the laptop, talking to the other side, they add headphones, producing V3 with charger, case and headphones. V2 and V3 are concurrent. A union merge gives charger, case and headphones: the case the customer deleted is back.

The root cause is that a set of items cannot tell "never added" from "added then removed". The fix is to record removals as data. An observed-remove set gives every add a unique tag, and a remove deletes only the tags it has seen. A concurrent add on another device carries a new tag, so it survives; the removed tag stays removed. This is a CRDT, and CRDT replication covers the general theory.

import uuid
from dataclasses import dataclass, field

@dataclass
class Cart:
    # adds:    tag -> (sku, qty)   every add gets a unique tag
    # removes: set of tags          a remove only kills the adds it has SEEN
    adds: dict = field(default_factory=dict)
    removes: set = field(default_factory=set)

    def add(self, sku, qty):                # local add: new total for this line
        self.set_qty(sku, self.items().get(sku, 0) + qty)

    def remove(self, sku):
        self.removes |= {t for t, (s, _) in self.adds.items() if s == sku}

    def set_qty(self, sku, qty):            # remove the tags we have seen, then one new tag
        self.remove(sku)
        if qty > 0:
            self.adds[uuid.uuid4().hex] = (sku, qty)

    def items(self):
        live = {}
        for t, (sku, qty) in self.adds.items():
            if t not in self.removes:
                live[sku] = max(live.get(sku, 0), qty)   # concurrent adds of one SKU: keep the larger
        return live

def merge(a: Cart, b: Cart) -> Cart:
    return Cart(adds={**a.adds, **b.adds}, removes=a.removes | b.removes)

Replaying the example: V2 records the case's tag as removed; V3 carries the same tag plus a new headphones tag. The merge unions both maps, the case's tag is in the remove set, and the result is charger and headphones, which is what the customer meant. Quantity conflicts need a policy too. Taking the maximum of concurrent quantities avoids doubling an order when the same add is replayed; summing would. Tombstone tags grow without bound, so compact them when all replicas have acknowledged a version, or keep a bounded window and accept rare resurrection for very old carts.

If your store has a single leader

Most teams today build carts on a managed database with a single writer per key and conditional writes, not on a leaderless Dynamo-style store. Concurrent writes then cannot silently diverge; instead they collide. Handle that with optimistic concurrency on a version number.

# Single-leader store (for example a key-value table with conditional writes):
# optimistic concurrency on a per-cart version number.
def add_item(cart_id, sku, qty, max_attempts=5):
    for _ in range(max_attempts):
        cart = store.get(cart_id) or {"version": 0, "items": {}}
        items = dict(cart["items"])
        items[sku] = min(items.get(sku, 0) + qty, MAX_QTY_PER_LINE)
        if len(items) > MAX_LINES:
            raise CartFull()
        ok = store.put_if(cart_id,
                          {"version": cart["version"] + 1, "items": items},
                          condition={"version": cart["version"]})   # compare-and-set
        if ok:
            return items
    raise Contention(cart_id)        # rare: two tabs hammering one cart

This gives you serialised, conflict-free carts at the cost of write availability: when the leader for a key is unavailable, the write fails and the client must retry. For multi-region deployments you are back to choosing between routing all writes for a cart to one home region or accepting concurrent writes and merging. The OR-set above works on either design; storing the cart as one document versus one item per row is a separate choice. One document makes reads a single lookup and the version check trivial. One row per item makes adds cheap and avoids rewriting a large cart, but the version check must cover the whole cart or you lose the per-line cap and the line-count limit.

Guest carts and the sign-in merge

Anonymous shoppers get a cart keyed by a random, unguessable token stored in a cookie. Give guest carts a time-to-live measured in weeks and refresh it on activity, so abandoned carts expire without a cleanup job.

When the guest signs in, two carts exist and must become one. The common rule is to add guest lines into the account cart, capping quantities per line, and to show the customer what changed. The merge must be idempotent: sign-in requests are retried, and a double merge doubles quantities. Tie it to an identifier from the sign-in request, as described in idempotency architecture, store that identifier in the same conditional write as the merged lines, and delete the guest cart only after the merge is durable.

def on_sign_in(customer_id, guest_token, merge_id):
    # merge_id comes from the sign-in request. It is written in the SAME conditional put as
    # the merged items, so a retry, a crash or two racing sign-ins cannot merge twice.
    guest = store.get(guest_token)
    while guest:
        cart = store.get(customer_id) or {"version": 0, "items": {}, "merges": []}
        if merge_id in cart["merges"]:
            break                                        # already merged
        items = dict(cart["items"])
        for sku, qty in guest["items"].items():
            items[sku] = min(items.get(sku, 0) + qty, MAX_QTY_PER_LINE)
        new = {"version": cart["version"] + 1, "items": items,
               "merges": (cart["merges"] + [merge_id])[-20:]}   # keep recent ids only
        if store.put_if(customer_id, new, condition={"version": cart["version"]}):
            break
    store.delete(guest_token)                    # only after the merge is durable
    return store.get(customer_id)

Reading the cart: composition with deadlines

Rendering the cart fans out to pricing, availability, delivery estimation and promotions. Each call gets a deadline well inside the page's latency budget, and each has a degraded answer: if delivery estimates time out, show the cart without them; if promotions time out, show list prices. The cart contents themselves must never be the part that is missing.

Store the price the customer saw when they added the item and compare it with the current price at render time. If it moved, say so on the page. This is both a trust feature and a correctness guard: the customer never discovers the change at the payment step. Cache catalogue data aggressively, but not stock or price used for the final decision, which is revalidated at checkout.

The checkout handoff

Checkout is where looseness ends. The protocol is:

  1. Read the cart with a strong read and record its version.
  2. Revalidate every line: current price, availability, purchase limits, whether the seller or offer still exists. Show differences before the customer confirms.
  3. Create a checkout session holding the snapshot and an idempotency key, so a double click or a retried request produces one order.
  4. Hand the snapshot to the order system, which reserves stock and charges payment. That half of the story is in designing an event-driven order system.
  5. When the order is accepted, clear the cart conditionally: if the version still matches the snapshot, empty it; otherwise remove only the ordered lines. An item added in another tab during checkout must survive.

Clearing the whole cart unconditionally is a classic bug: the customer adds an item on the phone while paying on the laptop, the order completes, and the new item vanishes.

Failure modes

FailureSymptomMitigation
Union merge without remove recordsDeleted items reappear after an outageOR-set with tags; alert on resurrection reports
Duplicate add replayedQuantity doublesClient request identifiers; max, not sum, for concurrent quantities
Guest merge retriedQuantities doubled at sign-inIdempotent merge keyed on a merge identifier
Unconditional cart clearItems added during checkout disappearClear ordered lines by version
Hot or huge cartsOne key dominates a partition; slow readsCap line count and quantity; rate-limit bots per cart
Dependency timeout on renderWhole cart page failsDeadlines and degraded answers per dependency
Price moved since addComplaints, charge disputesStore seen price, show change, revalidate at checkout

Operating the cart

Measure add-to-cart success rate and latency at the 99.9th percentile per region, because those are the promises. Track the rate of concurrent versions requiring a merge, the rate of conditional-write retries, and the size distribution of carts. A rising merge rate usually means a replication or network problem long before customers notice. Load-test with realistic traffic shapes: seasonal peaks are dominated by adds, and bot traffic concentrates on a few SKUs and sometimes on single carts. Treat cart deletion and expiry as data-protection features too; carts contain personal data and need retention rules.

Trade-offs

Leaderless, merge-on-read storage buys write availability during partitions at the cost of merge logic, tombstones and occasional surprising results. Single-leader storage with conditional writes buys simple, serialised carts at the cost of failed writes during failover and cross-region latency for multi-region customers. Storing prices in the cart simplifies rendering and makes it wrong; composing at read time costs fan-out and needs good timeouts. The cart can afford to be loose because checkout is strict; move validation out of checkout and every earlier shortcut becomes a correctness bug.

What to do next

  1. Write down whether your cart store is leaderless or single-leader, and what a customer sees when a write fails or two writes conflict.
  2. Model removals explicitly, with an OR-set or a version check, and add a test that removes on one device while adding on another during a simulated partition.
  3. Make adds, guest merges and checkout idempotent with request identifiers, and test each with a replayed request.
  4. Stop storing authoritative prices or stock in the cart; store the seen price only for change notices.
  5. Give every render dependency a deadline and a degraded answer, and verify it by injecting timeouts.
  6. Clear carts after an order by version, and add an end-to-end test that adds an item during checkout.
Key takeaway: The cart is a record of intent, not a reservation, so it can prioritise always accepting writes, as the Dynamo paper's cart did with sloppy quorums, hinted handoff and vector clocks. That availability creates conflicts, and a union merge resurrects deleted items; record removals with an OR-set or serialise with conditional writes. Keep prices and stock out of the stored cart, compose the page with deadlines, make adds and guest merges idempotent, and concentrate correctness at checkout: validate a versioned snapshot and clear the cart conditionally.