All 40 articles, sorted alphabetically
The CAP Theorem
The CAP Theorem explains why distributed systems cannot guarantee consistency, availability, and partition tolerance simultaneously. Understand the tr…
Read article →In Search of a Leader: Understanding Raft Consensus - AICassindra
In the world of distributed systems, getting a cluster of nodes to agree on something—like the order of log entries—is notoriously difficult.
Read article →API Gateway Patterns
Kong, Ambassador, AWS API Gateway. Auth (JWT validation), rate limits (per API key), routing (path-based, host-based, weighted), request/response tran…
Read article →Bulkhead Pattern
Isolate resources per feature. Prevent noisy neighbor.
Read article →Byzantine Fault Tolerance, in depth: the 3f+1 bound, PBFT's three phases, HotStuff's linear view change, and when BFT earns its cost
Byzantine fault tolerance from first principles: what the Byzantine fault model covers, why n must be at least 3f+1, partial synchrony and what it mea…
Read article →Circuit Breaker Pattern
Fail fast when dependency down. Prevent cascade failures.
Read article →CRDTs in Production, in depth: choosing data types, syncing with state vectors, storing and compacting history, evolving schemas and securing merges
What it takes to run conflict-free replicated data types in a real product: picking counter, register, set, map and sequence types by their surprising…
Read article →Anti-entropy -- healing divergence between replicas
Deep-dive on anti-entropy: the replica-divergence problem, background reconciliation, Merkle trees for efficient difference detection (O(1)/log(n)), r…
Read article →Bounded staleness architecture
Deep-dive on bounded-staleness consistency: version-based and time-based bounds, follower applied position and lag, read-path enforcement (serve, wait…
Read article →Causal consistency architecture
Deep-dive on causal consistency: the happened-before relation and session guarantees, dependency tracking with vector clocks, hold-until-ready applica…
Read article →Chain replication -- strong consistency with simple roles
Deep-dive on chain replication: the chain structure (head to tail), writes propagating down and committing at the tail, reads from the tail (committed…
Read article →Raft consensus architecture
Deep-dive on Raft consensus: roles, log replication, commit index, snapshots, joint-consensus membership, and the operational surface.
Read article →Raft Consensus, in depth: from the rules in the paper to a production node
Raft as you would implement and operate it: persistent state and fsync ordering, the vote and append handlers as code, why leaders only count replicas…
Read article →CRDT replication architecture
Deep-dive on CRDT replication: state-based vs op-based vs delta CRDTs, join-semilattice merges, OR-Sets and PN-counters, version vectors and dots, seq…
Read article →Distributed Hash Table Architecture in Depth
DHT architecture in depth: one id space for keys and nodes, virtual nodes, Chord fingers vs Kademlia k-buckets vs Pastry prefixes, join/leave converge…
Read article →Fencing tokens -- making distributed locks safe against pauses
Deep-dive on fencing tokens: the paused-lock-holder problem (GC/network/VM pauses past the lease), the monotonic fencing token issued per grant, resou…
Read article →Gray failure architecture
Deep-dive on gray failure: the degraded component whose shallow health check passes while real requests suffer, why differential observability is both…
Read article →Hinted handoff
Deep-dive on hinted handoff, the Dynamo-style mechanism that preserves write availability when a replica is temporarily down: a live custodian stores …
Read article →Hybrid Logical Clocks architecture
Deep-dive on Hybrid Logical Clocks: physical + logical combined, update rule, monotonicity, causality, bounded skew.
Read article →Distributed Leader Election Architecture in Depth
A 2500-word walkthrough of leader election in modern distributed systems: heartbeat, election timeout, PreVote, RequestVote, majority, split vote, and…
Read article →Leases -- time-bounded exclusive rights
Deep-dive on distributed leases: the safe-exclusive-access need, the time-bounded grant, expiry (auto-release), renewal (by the live holder), contrast…
Read article →Merkle tree architecture
Deep-dive on Merkle trees for replica reconciliation: leaf hashes over key ranges, parent hashes up to a single root, root-then-descend comparison, dr…
Read article →Operational transformation architecture
Deep-dive on OT for collaborative editing: the operation model, the transformation function and TP1/TP2, the Jupiter central-server model, client pend…
Read article →Transactional outbox architecture
Deep-dive on the transactional outbox pattern: how writing the business change and the event to publish in one local database transaction eliminates t…
Read article →Paxos
How classical Paxos works: proposers, acceptors, learners, and the two-phase protocol that achieves consensus.
Read article →Phi-accrual failure detectors
Deep-dive on phi-accrual failure detection: why fixed timeouts fail, the sampling window and gap-distribution estimator, how phi maps silence to log-s…
Read article →Quorum systems
Deep-dive on quorum-based replication: the R+W&amp…
Read article →Raft log replication
Deep-dive on Raft log replication: leader-based writes, log entries with term/index/command, AppendEntries replication, majority commit for fault tole…
Read article →Read repair and anti-entropy architecture
Deep-dive on the repair mechanisms behind leaderless eventual consistency: read repair on the coordinator&…
Read article →Saga Pattern, in depth: compensations, the pivot step, a durable orchestrator and the isolation you give up
How the saga pattern keeps multi-service business transactions consistent without distributed locks: compensatable, pivot and retriable steps, orchest…
Read article →Total order broadcast architecture
Deep-dive on total order (atomic) broadcast: the agreement, total-order, validity, and integrity guarantees, equivalence to consensus, leader-sequence…
Read article →Vector clock architecture, in depth: causal delivery, consistent cuts and debugging distributed traces
Vector clocks as event-level infrastructure: a clock layer between application and transport, causal broadcast with a hold-back queue, recovering cons…
Read article →Vector Clocks in Depth: Tracking Causality, Detecting Concurrent Writes and Keeping the Metadata Bounded
How vector clocks capture happens-before exactly: first principles, the update and comparison rules with Python code, a worked three-process example, …
Read article →Replication watermarks architecture
Deep-dive on watermarks in a replicated log: the high-water mark as the committed, safe-to-read boundary computed from the minimum offset replicated t…
Read article →ZooKeeper architecture, in depth: znodes, the ensemble, Zab, sessions, watches, consistency guarantees and how to run it
How Apache ZooKeeper works from the inside: the znode data model and node types, leaders, followers and observers, the Zab atomic broadcast and zxids,…
Read article →Distributed Transactions in Practice, in depth: choosing between co-location, XA, distributed SQL, sagas and idempotent retries
A practical guide to making changes atomic across machines and services: the decision ladder, how two-phase commit and XA fail in production, how Perc…
Read article →Failure Detectors and Phi Accrual, in depth: completeness, accuracy, the detector classes, and building timeout, phi and SWIM detectors that behave in production
How distributed failure detectors really work: why they can only suspect, completeness and accuracy, the Chandra-Toueg classes, detection time and mis…
Read article →What Jepsen Taught Us, in depth: histories, nemeses, checkers and the failure patterns that keep coming back
What a decade of Jepsen analyses teaches about distributed systems: how a Jepsen test is built, why indeterminate operations matter, how Knossos and E…
Read article →Quorum Systems Explained, in depth: the intersection property, majorities, weighted votes, grids, Flexible Paxos and Byzantine quorums
Quorum systems from first principles: why intersection is the only property that matters, majority and weighted voting, grid quorums and the load vers…
Read article →Raft Consensus Intuition, in depth: deriving every rule by breaking a simpler design
Build Raft from first principles: start with a primary and a backup, break it with partitions and crashes, and add one rule per failure until you reac…
Read article →