Jepsen is a testing library and a body of published analyses by Kyle Kingsbury. Since 2013 it has taken databases, queues and coordination services that make strong claims, injected network partitions, crashes, pauses and clock faults, and checked whether the system still did what its documentation promised. Very often it did not. Acknowledged writes disappeared, reads returned values that had already been overwritten, and transactions labelled serializable produced anomalies that serializability forbids.

This article is about what those results teach, not a list of who failed: how a test is built, why the history is the core idea, how checkers reach a verdict, which faults matter, and the few design mistakes behind most findings. A worked example walks through a history no correct register could produce.

Advertisement

The shape of a Jepsen test

A Jepsen test runs on a control node that reaches, by default, five database nodes named n1 to n5 over SSH. The test is a map of a few components. The :db component installs and starts the system under test. The :client translates abstract operations such as read, write or compare-and-set into real requests. The :generator decides which operations to issue and when. The :nemesis injects faults. The :checker analyses what happened afterwards.

The crucial design choice is that the database is a black box. Jepsen only records what clients saw: which operation each process invoked, when, and how it completed. That makes the method portable, and the findings hard to argue with, because the evidence is observations any user could have made.

A Jepsen test: one control node drives clients and faults, records a history, then checks itGeneratorinvoke ops, schedule faultsClientsone logical process eachNemesispartition, kill, pause, skewDatabase nodesn1 n2 n3 n4 n5Historyinvoke, ok, fail, infoCheckerKnossos or ElleVerdictvalid, invalid, unknownopsrequestsfaults via SSHcompletionsThe database is a black box: Jepsen only sees what clients observed, and when.Every claim it makes is a statement about that history, never about the source code.
The components of a Jepsen test. Faults and requests run concurrently; the checker only sees the recorded history.

The generator below follows the Jepsen tutorial: mixed register operations about 50 milliseconds apart, while the nemesis partitions the cluster into random halves for five seconds, heals for five, and repeats.

;; From the shape used in the Jepsen tutorial: a CAS register under partitions.
(defn r   [_ _] {:type :invoke, :f :read,  :value nil})
(defn w   [_ _] {:type :invoke, :f :write, :value (rand-int 5)})
(defn cas [_ _] {:type :invoke, :f :cas,   :value [(rand-int 5) (rand-int 5)]})

(defn register-test [opts]
  (merge tests/noop-test opts
    {:name      "register"
     :db        (db "v3.5.x")                 ; installs and starts the system on n1..n5
     :client    (Client. nil)                 ; maps ops to real requests
     :nemesis   (nemesis/partition-random-halves)
     :checker   (checker/compose
                  {:perf   (checker/perf)
                   :linear (checker/linearizable {:model     (model/cas-register)
                                                  :algorithm :linear})})
     :generator (->> (gen/mix [r w cas])
                     (gen/stagger 1/50)
                     (gen/nemesis
                       (cycle [(gen/sleep 5) {:type :info, :f :start}
                               (gen/sleep 5) {:type :info, :f :stop}]))
                     (gen/time-limit (:time-limit opts)))}))

Histories and the four completion types

Every operation appears in the history twice: once as an :invoke and once as a completion. There are three kinds of completion, and the distinction between them is the single most important idea in Jepsen.

CompletionMeaningWhat the checker may assume
:okThe operation definitely happened, with this result.It took effect at some instant between invoke and completion.
:failThe operation definitely did not happen.It had no effect and can be removed.
:infoThe outcome is unknown: a timeout, a dropped connection, a crash.It may have taken effect at any time after invoke, including after the test ended, or never.

A client that reports a timeout as a failure is lying, and Jepsen's analyses show why the lie matters. A write that timed out may still be sitting in a leader's log and commit later. If the client records :fail, the checker assumes the write never happened, and a later read that returns its value looks impossible. If the client records :ok to be optimistic, a genuinely lost write is hidden. Only :info is honest, and it is also what makes checking expensive, because an indeterminate operation stays concurrent with everything that follows it.

The same holds in application code: a timeout is an unknown outcome, not an error, which is why idempotency keys and read-back checks exist.

Advertisement

A worked example: a history no register could produce

Here is a small history for a single register that starts at 0. Process 0 writes 3 and times out. Process 1 then reads 3. Later, process 2 reads 0.

{:process 0, :type :invoke, :f :write, :value 3,   :time 0}
{:process 0, :type :info,   :f :write, :value 3,   :time 10, :error :timeout}
{:process 1, :type :invoke, :f :read,  :value nil, :time 12}
{:process 1, :type :ok,     :f :read,  :value 3,   :time 13}
{:process 2, :type :invoke, :f :read,  :value nil, :time 20}
{:process 2, :type :ok,     :f :read,  :value 0,   :time 21}

Is this linearizable? Linearizability requires a single total order of operations, consistent with each operation's real-time interval, in which every read returns the latest write. The write is indeterminate, so it could have taken effect at any point after time 0. Process 1's read of 3 forces it to have taken effect before time 13. Process 2's read starts at 20, after process 1's read finished, so it must come later in the order, and there is no write of 0 to explain it. No legal order exists. The register went backwards, which is the signature of a stale read from a node that did not know it had been superseded.

A checker does this reasoning by search. The sketch below is the brute-force idea: pick an operation that could go next without violating real time, apply it to a model of the register, and backtrack if the model rejects it. Indeterminate operations may be omitted, because they may never have happened.

INF = float("inf")

def step(state, op):
    """Apply one op to a register model; return (legal, new_state)."""
    if op["f"] == "write":
        return True, op["value"]
    if op["f"] == "read":
        return op["value"] == state, state
    if op["f"] == "cas":
        old, new = op["value"]
        return state == old, (new if state == old else state)

def linearizable(ops, state=0):
    """Search for a legal order that respects real time. :info ops have
    end = INF and may be omitted, since they may never have taken effect."""
    if all(o["end"] == INF for o in ops):
        return True
    for o in ops:
        # o may go next only if no other pending op finished before o began
        if any(q["end"] < o["start"] for q in ops if q is not o):
            continue
        legal, nxt = step(state, o)
        if legal and linearizable([q for q in ops if q is not o], nxt):
            return True
    return False

history = [
    {"f": "write", "value": 3, "start": 0,  "end": INF},   # timed out: :info
    {"f": "read",  "value": 3, "start": 12, "end": 13},
    {"f": "read",  "value": 0, "start": 20, "end": 21},
]
print(linearizable(history))   # False: the register went from 3 back to 0

Checkers: Knossos and Elle

Checking linearizability of a general history is NP-complete, a result from Gibbons and Korach in 1997. Jepsen's original checker, Knossos, searches for a legal order with pruning and memoisation. Its cost grows sharply with concurrency and with indeterminate operations, which is why tests use small value ranges, many independent keys and short runs.

Transactional systems needed something else, and Elle, published by Kingsbury and Peter Alvaro in 2020, supplies it. Elle chooses workloads whose histories reveal the order of versions directly. In the list-append workload each transaction appends unique values to lists and reads whole lists, so any read shows exactly which appends preceded it. From that Elle infers write-write, write-read and read-write dependencies between transactions, builds a dependency graph in the style of Adya's isolation formalism, and looks for cycles. A cycle is a concrete anomaly with a name: G0 for write cycles, G1c for cycles of writes and reads, G-single for read skew, G2 for anti-dependency cycles that serializability forbids. It also flags G1a (reading aborted data) and G1b (reading intermediate data).

Elle scales to large histories and explains each verdict with a concrete cycle of transactions a vendor can reproduce.

Nemeses: the faults that matter

FaultHow it is injectedWhat it tends to expose
Network partitioniptables rules: random halves, isolated nodes, a majorities ring, a bridgeSplit brain, lost acknowledged writes, stale reads from deposed leaders
Process crashKill the database process and restart itWrites acknowledged before they were durable, recovery bugs
Process pauseSIGSTOP then SIGCONT, which also mimics a long GC pauseLeases trusted after they expired, zombie leaders
Clock faultBumping or strobing the system clock on some nodesTimestamp ordering bugs, lease violations, last-write-wins data loss
Disk faultLosing writes that were never fsynced, corrupting filesMissing fsync calls, torn logs, bad checksums
Membership changeAdding and removing nodes during the testReconfiguration bugs, which are among the hardest consensus code

These faults are ordinary: switches misbehave, JVMs pause for garbage collection, VMs migrate, clocks jump. A system correct only without them is correct only in the lab.

The recurring patterns

Read many analyses and the same few mistakes recur. The examples name the analysis where a pattern appeared; reports are version-specific and many issues were later fixed, so read them as illustrations, not statements about current releases.

  • Acknowledgement is not durability across failover. Kafka 0.8 beta (2013) could lose acknowledged writes when the in-sync replica set shrank to the leader alone and a lagging replica was then elected. Redis 2.6.13 with Sentinel (2013) replicates asynchronously, so a failover discards whatever the old primary had acknowledged but not shipped. MongoDB 2.4.3 (2013) rolled back acknowledged writes after partitions, even with a majority write concern, because of a bug that acknowledged during a partition. The client heard yes before a majority durably had the data.
  • Leaders that do not know they have been deposed. The 2014 analysis of etcd 0.4.1 and Consul showed stale reads because a leader answered reads locally, and a leader on the minority side of a partition can keep doing that after the majority elects someone else. The fix is to confirm leadership with a quorum before answering, as Raft's read index does.
  • Home-grown consensus. Elasticsearch 1.1.0 (2014) and 1.5.0 (2015) lost writes under partitions when its cluster coordination could elect two primaries. Ad hoc consensus rarely survives a nemesis.
  • Trusting clocks. Cassandra 2.0.0 (2013) resolves conflicts by timestamp, so skewed clocks silently discard newer writes. The CockroachDB beta analysis (2017) confirmed that its guarantees depend on a maximum clock offset, and showed what happens when that bound is exceeded.
  • Isolation levels that are not what they say. Elle found a G2-item anomaly in PostgreSQL 12.3 (2020) under serializable isolation, which was then fixed. MongoDB 4.2.6 (2020) showed read skew, cyclic information flow and duplicate writes in transactions even at snapshot read concern and majority write concern. The MySQL 8.0.34 analysis (2023) found that its repeatable read level permits anomalies that the formal definition of repeatable read excludes.
  • Defaults weaker than the documentation. A recurring theme is that a system could be configured safely, but the defaults, or the wording of the documentation, implied more than the defaults delivered.

There is a positive example too: ZooKeeper 3.4.5 (2013) passed. Correctness under faults comes from a proven protocol, carefully implemented and actually tested with faults.

Applying the lessons to a system you operate

Most of the fixes are configuration and client behaviour, so you can act on the analyses without writing a test.

  • Ask for majority acknowledgement. In Kafka: acks=all, min.insync.replicas of at least 2, and unclean.leader.election.enable=false. In MongoDB: a majority write concern and a matching read concern. This is necessary, not sufficient, as the MongoDB results show.
  • Know which reads are linearizable. Many systems offer a fast read that can be stale and a slower quorum or leader read. Choose per call site, deliberately.
  • Treat timeouts as unknown. Use client-generated idempotency keys, retry only idempotent operations, and read back when you must know the outcome.
  • Monitor clocks wherever guarantees depend on bounded skew, and alert well before the bound.

Testing your own system the Jepsen way

The method transfers without the library. Record every client operation with invoke time, completion time and an honest completion type, run a workload on a small cluster while a fault injector partitions, kills, pauses and skews nodes, and check the history against a model. A set workload is the cheapest start: add unique elements, read the whole set after healing, and require every acknowledged add to be present. Keep runs short, repeat them many times, and save the smallest failing history as the bug report.

Limits and trade-offs of the method

  • It finds bugs; it does not prove their absence. Five nodes and a few minutes cannot reach every state a production cluster reaches over months.
  • Checker cost. Linearizability checking is exponential in the worst case, which constrains workload design.
  • Black-box reach. Deterministic simulation reaches rare interleavings more reliably, and TLA+ specifications catch design flaws before code exists. Mature teams use all three.

For the protocols behind these findings, see Raft consensus in depth and Paxos. For what a quorum acknowledgement actually buys you, read quorum systems, and for detecting concurrent writes without trusting clocks, vector clocks.

What to do next

  1. Find the Jepsen analysis for each database and queue you run, note the version it covered, and check which findings your version still has.
  2. Audit client settings for write acknowledgement and read consistency, and change any call site that relies on defaults it has not checked.
  3. Change client code so a timeout is handled as an unknown outcome, and add idempotency keys to every retried write.
  4. Add clock-skew monitoring wherever correctness depends on timestamps or leases.
  5. For a service you build, write a set workload with a partition and kill nemesis, record histories with invoke, ok, fail and info, and run it in CI.
  6. Keep the smallest failing history from every bug you find, and turn it into a regression test.
Key takeaway: Jepsen's lasting contribution is a method: treat the system as a black box, record what clients observed with honest completion types, inject the faults that happen in real deployments, and check the history against a formal model. The recurring findings are few. Acknowledgements came before a majority had the data, leaders answered reads after being deposed, home-grown consensus split its brain, clocks were trusted, and isolation levels were weaker than their names. Configure for quorum acknowledgement, choose read consistency deliberately, treat timeouts as unknown, and test your own services with faults and a checker.