The whole model in one line

Every kernel does some arithmetic and moves some bytes. Call the work W (FLOPs) and the traffic Q (bytes crossing to the memory level you care about). Their ratio is the arithmetic intensity:

I = W / Q          [FLOP per byte]

P(I) = min( P_peak , BW · I )   [attainable FLOP/s]

That is the entire model. The machine can never issue more than P_peak FLOP/s nor feed itself faster than BW bytes per second, so a kernel of intensity I cannot exceed BW · I FLOP/s however good your code is. Plot P(I) on log-log axes and you get a slope-1 line flattening into a horizontal ceiling: a roof. The model, from Williams, Waterman and Patterson, is deliberately crude; its value is being an upper bound you can compute on paper, and upper bounds stop you optimizing the wrong thing.

Advertisement

The ridge point splits the world in two

The bend where slope meets ceiling is the ridge point. Left of it a kernel is bandwidth-bound: it finishes when its bytes arrive, and every FLOP beyond is idle capacity. Right of it, compute-bound. Two rules follow, and they are the only two you need: cutting FLOPs is worthless left of the ridge, cutting bytes right of it. Most wasted LLM optimization work is a right-of-ridge fix on a left-of-ridge kernel. Setting the two terms equal locates it:

I_ridge = P_peak / BW    (machine balance)

Our reference machine throughout: 8 cores at 3.5 GHz, AVX2 (8 fp32 lanes), two FMA ports. An FMA is 2 FLOPs per lane, so 8·2·2 = 32 FLOP/cycle/core and P_peak = 8·3.5e9·32 ≈ 896 GFLOP/s. Dual-channel DDR5-5600 gives 2·5600e6·8 = 89.6 GB/s, so I_ridge = 896/89.6 = 10 FLOP/byte. A datacenter accelerator (1000 TFLOP/s bf16, 3.3 TB/s HBM) sits near 300, thirty times harder to feed — the quiet argument for CPU small-model serving.

Advertisement

Building the two ceilings honestly

A roofline is only as good as its roof, and spec-sheet numbers are usually not the numbers you can reach. Peak FLOP/s assumes every core issues a full-width fused multiply-add every cycle at nominal clock — but sustained AVX-heavy code clocks down, and a kernel that cannot use FMA or fill the vector width loses a large multiple immediately. Peak bandwidth is likewise theoretical; a STREAM triad on real hardware lands at 70–85% of it.

So measure both: a small cache-resident FMA loop for the compute ceiling you can actually hit, STREAM or any large-array copy for the bandwidth you can actually sustain. A roofline from measured ceilings is a diagnostic; one from marketing numbers just tells everybody their kernel is at 8% of peak. The ridge moves with the measured numbers too: softer ceilings push it slightly left.

Worked example: a prefill GEMM

Take a small model with d = 2048, FFN width 8192, fp16, prefilling a 512-token prompt. The up-projection is C[512, 8192] = A[512, 2048] · B[2048, 8192].

FLOPs  W = 2·M·N·K = 2 · 512 · 8192 · 2048 = 17.2 GFLOP
bytes  Q = (512·2048 + 2048·8192 + 512·8192) · 2 B = 44.0 MB
I      = 17.2e9 / 44.0e6 ≈ 390 FLOP/byte

An intensity of 390 sits 39× to the right of our ridge of 10. The bandwidth roof there would be 89.6 · 390 ≈ 35 TFLOP/s, absurdly higher than the machine can compute, so the binding constraint is unambiguously P_peak. Prefill is compute-bound with room to spare, and its levers are vectorization, FMA utilization, thread scaling and tiling for cache reuse — not shrinking weights. Cutting bytes here buys you nothing.

Worked example: a decode-step matvec

Now generate one token. The same layer becomes M = 1: a matrix-vector product against the identical [2048, 8192] weight matrix.

FLOPs  W = 2 · 1 · 8192 · 2048 = 33.6 MFLOP
bytes  Q ≈ 2048·8192·2 B (weights dominate)  = 33.6 MB
I      = 33.6e6 / 33.6e6 ≈ 1.0 FLOP/byte

The result is not an accident of these dimensions but structural: a matvec touches every weight once and does 2 FLOPs with it, so in fp16 (2 bytes per weight) intensity is pinned at 2 FLOP / 2 B = 1 whatever the model size. Intensity 1 against a ridge of 10 means the attainable rate is min(896, 89.6 · 1) = 89.6 GFLOP/s: exactly 10% of peak. Ninety percent of the arithmetic hardware is structurally unreachable during decode, and no kernel tuning changes it — the limit is the weight sweep itself.

The same machine, both workloads, one plot

Drawn on the reference machine, the two land on opposite sides of the bend. Decode sits on the slanted roof, already optimal in the only sense that matters; prefill sits below the flat roof, which is the shape of a real optimization opportunity.

1010010000.251416642561024no-SIMD ceiling — 56 GFLOP/sridge: I = 10 FLOP/byte896 GFLOP/s ÷ 89.6 GB/scompute roof — 896 GFLOP/sbandwidth roof — 89.6 GB/sdecode matvec: I ≈ 189.6 GFLOP/s = 10% of peakprefill GEMM: I ≈ 390measured 540 — 60% of roofarithmetic intensity — FLOP per byte of DRAM traffic (log)attainable GFLOP/s (log)
One machine, two workloads. Prefill sits far right under the flat compute roof; decode sits on the slope, capped at a tenth of peak.

Reading the gap, and the roofs below the roof

The vertical gap from a plotted point to the roof above it is your headroom, in three flavours. No gap means you are done — only a horizontal move helps. A modest gap under the flat roof is classic kernel work: the prefill GEMM above measures 540 GFLOP/s against an 896 roof, 60% of peak — poor vectorization or thread imbalance. A large gap under the slanted roof is the interesting one — you are neither computing nor streaming at capacity, so the limit is elsewhere: dependent-load latency, NUMA traffic, or strided access wasting each cache line.

Sub-ceilings turn that diagnosis into a checklist: the peak without FMA (half), for scalar code at one FMA per cycle (8·3.5e9·2 = 56 GFLOP/s, sixteen times down), on one thread (an eighth). A point on the no-SIMD line is not a memory problem — the compiler failed to vectorize your inner loop. Precision adds its own family of roofs, so always say which one a plot uses.