A message is a block of tokens, read twice

Start with the unit. When agent A sends agent B a message, that message costs tokens at least twice: once when A generates it (decoding, one token at a time) and again when B reads it (prefill, ingesting it into B’s context). If the message is broadcast to several agents, the read cost is paid once per recipient. So the natural cost model for a message of m tokens sent to r recipients is:

cost(message) = m (generate)  +  r × m (each recipient prefills)

The generation term is fixed by whoever speaks; the r × m read term is where topology enters, because r is decided entirely by who is wired to hear whom. And crucially, prefill is not a one-time charge: on a stateless chat API every agent re-sends its whole context every turn, so a message that lands in a shared transcript is re-read on every subsequent round, not just the round it arrived. That re-reading, compounded over rounds, is what turns a modest conversation into a runaway bill.

Advertisement

Shared state versus private state

The first design fork is whether agents share a single context or each keep their own. A shared or ‘blackboard’ design gives every agent the full transcript: everyone sees everything, so coordination is easy and nobody misses a fact. A private design gives each agent only the slice it needs — its own task and whatever was explicitly forwarded to it.

The tradeoff is pure token economics. Shared state means every message enters a transcript that every agent re-ingests each turn, so redundant reading grows with both the number of agents and the number of messages. Private state removes that redundancy — an agent never pays to read a conversation it is not part of — but it risks divergence: two agents can now hold contradictory pictures of the world and never notice. Put bluntly, shared state buys coherence with tokens; private state saves tokens and risks incoherence. Most robust systems land in between: private working context plus a small, deliberately-shared summary of the facts everyone must agree on.

Advertisement

Topology decides the fan-out

The communication graph — who can talk to whom — is the single biggest lever on cost, because it fixes the recipient count r for every message. Three shapes cover most real systems:

TopologyEdgesWho hears a message
Star (orchestrator + workers)N − 1Only the hub, or one worker the hub picks
Chain / pipelineN − 1Only the next stage
Full graph (mesh)N(N − 1)/2Every other agent

Edge count is the headline. A star and a chain both connect N agents with only N − 1 links, so communication scales linearly. A full mesh has O(N^2) links, and if every agent broadcasts to every other each round, the token traffic scales like N^2 too. The jump from linear to quadratic is not a rounding error — it is the difference between a team that scales to dozens of agents and one that chokes at five.

The quadratic cost of a full mesh

Make the mesh cost concrete. Suppose N agents each speak once per round with a message of m tokens, and every message is broadcast to the other N − 1 agents. The read (prefill) traffic for a single round is:

reads_per_round = N × (N − 1) × m  ≈  N^2 · m

Now add the compounding. On a shared transcript, round t forces each of the N agents to re-read everything said in rounds 1…t−1. Summing the growing transcript over T rounds gives:

total_reads ≈ N^2 · m · Σ_{t=1}^{T} (t − 1)
            = N^2 · m · T(T − 1)/2   =   O(N^2 · m · T^2)

Quadratic in agents and quadratic in rounds. This double-quadratic is the signature failure mode of naive ‘let all the agents debate’ designs: doubling the team quadruples the fan-out, and letting the debate run twice as long quadruples the re-reading. It looks fine in a three-agent demo and detonates in production.

A worked token budget

Plug in numbers. Take N = 5 agents, messages of m = 300 tokens, running for T = 6 rounds. In a full-mesh shared transcript, the cumulative read cost is:

total_reads = N^2 · m · T(T−1)/2
            = 25 × 300 × (6·5/2)
            = 25 × 300 × 15  =  112,500 tokens

Generation is almost a rounding error next to that: N · m · T = 5 × 300 × 6 = 9,000 tokens. So roughly 92% of the compute goes to agents re-reading each other, not to producing anything new. Now switch to a star where the orchestrator compresses each round into a d = 150-token digest and workers see only that digest plus their ~200-token task. Worker reads become N · T · (d + task) = 5 × 6 × 350 = 10,500; the hub reads N · m replies per round, 5 × 300 × 6 = 9,000. Total ≈ 19,500 tokens — the same six rounds of collaboration for under a fifth of the cost.