Why architecture matters here

A transactional engine turns a promise — these statements all happen or none of them do, and nobody observes a half-finished version — into a physical arrangement of a log file, a page cache and a lock table. This article walks that arrangement: what each ACID letter costs in mechanism, why the write-ahead log's ordering rule is durability, how ARIES-style recovery reconstructs a crashed instance, what the three families of concurrency control charge you, and which anomaly each isolation level does and does not permit — the part most often stated wrongly.

ACID as the engine implements it

The four letters are delivered by three mechanisms, and knowing which is which tells you what to tune.

Atomicity is undo. The engine must be able to erase a partial transaction, so it must remember the prior state of everything it changed — in a rollback segment (Oracle, InnoDB) or interleaved with redo information in one log stream (ARIES-style systems).

Consistency is mostly not the database's job. Declared constraints, foreign keys and triggers are enforced; an application invariant such as "these two balances always sum to 5000" is invisible to the engine and holds only if isolation is strong enough or the code takes explicit locks. Most write-skew bugs are consistency violations nothing promised to catch.

Isolation is concurrency control — locking, versioning, validation.

Durability is the log and one flush. Nothing is durable because it was written to a data page; it is durable because a commit record reached stable storage.

Advertisement
Advertisement

The write-ahead log: ordering is the guarantee

Every change generates a log record stamped with a monotonically increasing LSN and appended to a sequential file. The record is physiological — physical about which page it touches, logical about what it did inside that page ("insert this tuple into slot 4 of page 118" rather than a byte image), which keeps records small while keeping redo cheap. Each data page carries in its header the LSN of the last record applied to it, the pageLSN. Two ordering rules then do all the work.

The write-ahead rule. The log record describing a change to page P must be on stable storage before P itself is written out. Without it a crash can leave a modified page on disk with no record of how to undo it.

The commit rule. All of a transaction's records, up to and including its commit record, must be flushed before the client is told the commit succeeded. Without it you acknowledge work you cannot reconstruct.

Everything else about the WAL is a performance argument: the log is one sequential append shared by every session, while the pages it protects are scattered random writes that can be deferred, coalesced, and written once for many updates.

The architecture: every layer explained

Read the diagram as three layers. On top, the transaction manager tracks each transaction's state, its snapshot and its position in the log, while the requested isolation level decides which conflicts it must police. In the middle, concurrency control — version visibility and the lock manager together — decides who sees what and who waits. Underneath, the write-ahead log, the undo information and the recovery passes turn all of it into something that survives power loss, with vacuum reclaiming versions nobody can see any more and two-phase commit stretching atomicity across nodes. Each piece is unpacked below.

ClientBEGIN + statements + COMMITTransaction Managerstate + lock + snapshotIsolation LevelREAD COMMITTED / SERIALIZABLEMVCCrow versions + visibilityLock Managerrow / table / predicateWAL / Redo Logwrite-ahead durabilityRollback Segmentundo before commitVacuum / Autovacuumclean tombstoned rows2PC for distributedcoordinator + participantsRecoveryreplay WAL after crashPostgres, MySQL InnoDB, Oracle, Aurora, CockroachDB use MVCC + WAL
Transactional DB architecture: transactions with MVCC + locks + WAL/redo + rollback + isolation levels + 2PC for distributed + recovery.

Steal, no-force, and the debt they create

Two buffer-pool policies decide what recovery has to do. Steal lets the pool evict a dirty page belonging to an uncommitted transaction, putting uncommitted data on disk — which is what allows a transaction larger than memory, and exactly why undo information must exist. No-force means commit flushes only the log, not the transaction's data pages — which turns many random writes per commit into one sequential one, and exactly why redo information must exist.

PolicyUndoRedoRun-time cost
no-steal / forcenonopin every dirty page to commit, then flush them all
steal / forceyesnorandom write storm on every commit
no-steal / no-forcenoyestransaction size bounded by the pool
steal / no-forceyesyescheapest at run time, most work in recovery

Every serious engine picks steal/no-force and pays with a real recovery algorithm. Pool mechanics themselves — pins, latches, replacement — are covered in buffer pool architecture.

ARIES recovery: analysis, redo, undo

Checkpoints keep restart bounded. A fuzzy checkpoint does not flush the pool; it records the dirty page table (each dirty page with the recLSN of the oldest record that dirtied it) and the transaction table (each live transaction with its latest LSN). Restart then makes three passes.

Analysis scans forward from the checkpoint, rebuilding both tables as of the crash: which transactions were still live (the losers) and which pages may be dirty.

Redo starts at the smallest recLSN in the dirty page table and repeats history, reapplying every logged change — including changes made by losers. For each record it compares the record's LSN with the page's pageLSN and skips the page when the change is already there, which makes redo idempotent. Repeating history reconstructs the exact state at the crash, giving undo a well-defined starting point.

Undo rolls the losers back in reverse LSN order. Each undone action writes a compensation log record carrying an undoNextLSN pointer to the next record still owing a rollback. CLRs are redo-only and never themselves undone, so a crash during recovery resumes where the pointers left off rather than undoing the undo.