What multi-agent actually means here
A multi-agent LLM system is any setup where more than one model instance — or one model wearing several roles across turns — contributes to a single result through structured interaction. Agents share state, exchange messages, and their outputs are combined by some protocol: voting, a manager that routes subtasks, a critic that revises a proposer.
The unit of analysis is not the model but the interaction graph: who talks to whom, in what order, and how contributions are aggregated. Two systems running the identical base model behave completely differently when their graphs differ — a debate ring, a pipeline, a star, a fully connected mesh. Almost every interesting property is a property of that graph and its aggregation rule, not of the weights — which is why a theory of these systems can say useful things without opening the model at all.
The central question: committee versus soloist
The one question that governs everything is deceptively simple: does adding agents help, and if so, why? A single call has a fixed competence — some probability it lands the answer, some distribution over its mistakes. Adding agents helps only if the group accesses something the soloist cannot: independent samples of the answer, specialization into subtasks each agent does better, or an external check that catches errors the generator is blind to.
Notice what is not on that list: raw model quality. Multi-agent structure does not make any individual agent smarter — it only rearranges how limited, fallible agents are composed. So the honest framing is not ‘more agents are better’ but ‘when does composition extract value a single pass leaves on the table?’ Each answer below also marks when it extracts nothing and you are just paying more tokens.
The decomposition-versus-overhead tradeoff
Every multi-agent design lives on one axis. One side is the gain from decomposition: splitting a hard task into easier pieces, or sampling the answer several times so errors cancel. The other is the overhead of coordination: the extra tokens, latency, and failure modes that talking introduces. Net value is the difference.
Make the cost concrete. A single answer costs one generation. A system with k agents running r rounds, each round passing along a transcript that grows with the conversation, costs on the order of:
cost ≈ Σ_{t=1}^{r} k · (C_in(t) + C_out)
where C_in(t) grows with the shared transcript so farBecause C_in(t) swells as every agent appends to the transcript, a chatty r-round system grows faster than linearly — closer to O(k · r^2) tokens for fully shared history. Decomposition gains, by contrast, are bounded and hit diminishing returns after a handful of agents. The two curves cross, and past that crossing every added agent makes the answer worse per dollar.
A worked cost example
Suppose one agent solves a task correctly 70% of the time and you convene five that vote. With perfectly independent errors, majority vote of five raises accuracy to roughly 84% — a real gain. But five calls cost about 5× the tokens, and letting them discuss for three rounds first can push the bill past 10–15× the single call.
Now the money question: is 70% to 84% worth 10× the cost on this task? For a high-stakes contract clause, obviously yes. For classifying a support ticket, obviously no — ship the single call and spend the saved budget elsewhere. The theory does not pick for you; it makes the tradeoff legible. The common mistake is running the 10× committee on the support-ticket task because it ‘felt robust,’ never pricing the two curves against each other.
A taxonomy: eight problems under one roof
The concrete pieces in this series are not eight unrelated tricks — they are the components every multi-agent system assembles from, in three layers. Substrate is how agents exchange and retain information: communication protocols and shared memory. Interaction pattern is what they do with that channel: debate, negotiation, coordination, and reflection. System-level behavior is what results: multi-step planning and the emergence of group capabilities no single agent was given.
Read that way, the taxonomy is a build order: choose a substrate, pick an interaction pattern suited to the task, and the system-level behavior is the consequence you steer — not a fourth thing you bolt on.
Error correction versus error amplification
The single most important dynamic in multi-agent systems is which direction errors flow. In error-correcting mode, a mistake by one agent is caught and fixed by another: a critic flags a hallucinated citation, a majority outvotes an outlier, a verifier rejects a bad plan. The group is more reliable than any member. In error-amplifying mode, one agent’s mistake propagates and is reinforced: agents defer to a confident-but-wrong peer, an early error poisons the shared transcript, and the group locks onto it together.
The same architecture can do either, and which one you get is not decided by the prompt’s good intentions but by structural properties — independence, a genuine check, and how aggregation weights dissent. Debate and reflection try to engineer the correcting regime; sycophancy and groupthink are the amplifying regime showing up uninvited. The next two sections make the boundary quantitative.
The math of voting: independence is everything
Take the cleanest case, majority vote over k agents that each answer correctly with probability p. If their errors are independent, the probability the majority is right is the tail of a binomial, and for p > 0.5 it climbs toward 1 as k grows — the Condorcet jury theorem:
P(majority correct) = Σ_{i=⌈k/2⌉}^{k} C(k,i) · p^i · (1-p)^(k-i)
p = 0.7, k = 1 → 0.70
p = 0.7, k = 5 → 0.84
p = 0.7, k = 15 → 0.95This is the entire optimistic case for scaling agents, and its fine print is brutal: p > 0.5 and independence. If each agent is worse than a coin flip, voting drives accuracy toward zero — you amplify the error. And independence is exactly what shared context quietly destroys, which is the next section.