A tool bomb is an input that makes an agent do far more tool work than the input is worth. A web page that says to fetch each of its two hundred links, a ticket that asks a coding agent to rerun the tests until they pass, or a tool that returns forty thousand tokens where four hundred were expected: each turns a few kilobytes of text into hundreds of calls and a bill that arrives after the damage.
The name borrows from the zip bomb, a small archive that expands into terabytes. Agents multiply: one plan step can request many tool calls, a sub-agent can spawn its own, a failed call can be retried at three layers, and every tool output is re-sent to the model on every later turn. This page is about that multiplication. Request-level admission control and serving limits are covered in LLM Denial of Service, in depth, and misuse of a single legitimate tool in Tool Abuse, in depth. Here the subject is the call graph an agent builds for itself, and the budget tree that keeps it bounded; the tool server side is in MCP rate limiting.
First principles: the amplification factor
Security people measure amplification attacks with one ratio: the work the victim performs divided by the work the attacker spends. For agents the attacker spends some tokens of injected or submitted text; the victim spends tool calls, model tokens, external API fees, compute in sandboxes and the time of the people waiting. If 2 KB of planted text causes 300 tool calls and 6 million input tokens, the amplification factor is enormous, and it does not matter that no single call was malicious.
Three properties make agents unusually good amplifiers. First, the model decides the shape of the call graph at run time, so the developer never wrote the loop that runs away. Second, instructions and data share one channel, so a document can request work, as described in Indirect Prompt Injection, in depth. Third, costs compound rather than add: context re-sending makes token spend grow with the square of the number of turns, and nested delegation makes call counts grow exponentially with depth.
Many tool bombs are accidents, such as a flaky tool whose errors the model keeps trying to fix. The defence is the same either way, which helps, because you cannot tell intent from a running call graph.
The six shapes of amplification
Almost every incident fits one of six shapes, and each needs a slightly different limit.
| Shape | How it multiplies | Typical trigger | Limit that stops it |
|---|---|---|---|
| Fan-out | One step requests k parallel calls | List of URLs, IDs or files in tool output | Per-step call cap, per-run call budget |
| Recursion and delegation | Sub-agents spawn sub-agents; crawl follows links | Task split into subtasks, each split again | Depth limit and carved child budgets |
| Loops | Same or alternating calls repeat | Tool error the model tries to fix; goal never satisfied | Call fingerprinting, no-progress detector |
| Retry storms | Retries at agent, SDK and HTTP layers multiply | Upstream timeout or 429 | One retry layer, retries charged to the budget |
| Output inflation | Large output re-sent on every later turn | Tool returns full page, log or table | Output byte cap with an out-of-band handle |
| Expensive single call | One call costs dollars or minutes | Paid API, full table scan, VM creation, long job | Per-tool price ceiling and approval above it |
Shapes combine. The worst incidents are usually fan-out inside a loop, or delegation where every child fetches large documents. Retry storms are the most surprising because each layer looks reasonable alone: an agent framework that retries a failed step three times, a tool SDK that retries three times and an HTTP client that retries three times give twenty-seven attempts per logical call, all against an upstream that is already failing.
Why output size is the quiet one
Most agent loops append every tool result to the transcript and send the whole transcript back to the model on each turn. A tool output of o tokens therefore costs o tokens once to produce and o tokens again on every later turn. Over n turns the input tokens grow roughly as n squared times o over two. The arithmetic is easy to check:
def resend_cost(turns, tool_output_tokens, base_prompt=3_000):
# Every turn re-sends the whole transcript; each turn appends one tool output.
total = 0
context = base_prompt
for _ in range(turns):
total += context # input tokens billed this turn
context += tool_output_tokens # output joins the transcript
return total
print(resend_cost(20, 500)) # 155,000 input tokens: small outputs
print(resend_cost(20, 40_000)) # 7,660,000 input tokens: one fat tool, same 20 turnsThe same twenty-turn task costs about fifty times more when one tool starts returning full documents. Nothing failed and no call limit fired. Output inflation also crowds the context window, which degrades later decisions and makes loops more likely. That is why an output cap belongs at the tool boundary: by the time the window is full you have already paid for every turn that led there.
Watch for tools that expand data, too: a file tool that decompresses archives or an HTML tool that inlines frames turns small fetched bytes into large context bytes. Cap the bytes that enter the context, not only the bytes fetched.
The architecture: one ledger, carved budgets
The control that works is a budget ledger enforced at the tool gateway, the one place that sees every call regardless of which planner, prompt or sub-agent issued it. The model is told its limits, which helps it plan, but the gateway is what enforces them. Prompts are advice; the ledger is a lock.
Four rules make the ledger hard to escape. Budgets cover calls, tokens, money, wall clock, depth and calls per step at once, because each shape exhausts a different one. Child budgets are carved, never granted, so the tree total cannot exceed the root. Costs are reserved at the worst-case price before a call and settled after, so concurrent calls cannot overspend. And exhaustion is a normal result: the model sees a structured budget error and can summarise partial findings.
The ledger in code
The sketch below is framework-neutral. It keeps the counters behind a lock because fan-out usually runs calls in parallel, and it fingerprints each call so identical repeats trip a loop detector.
import hashlib, json, threading, time
class BudgetExceeded(Exception):
pass
class Budget:
def __init__(self, calls, tokens, dollars, seconds, depth=0, max_depth=2, parent=None):
self.left = {"calls": calls, "tokens": tokens, "dollars": dollars}
self.deadline = time.monotonic() + seconds
self.depth, self.max_depth, self.parent = depth, max_depth, parent
self.lock = threading.Lock()
self.seen = {} # call fingerprint -> count
def carve(self, calls, tokens, dollars, seconds):
# A child budget is taken OUT of this one, so the tree total can never grow.
if self.depth + 1 > self.max_depth:
raise BudgetExceeded("sub-agent depth limit")
self.charge(calls=calls, tokens=tokens, dollars=dollars)
remaining = max(0.0, self.deadline - time.monotonic())
return Budget(calls, tokens, dollars, min(seconds, remaining),
self.depth + 1, self.max_depth, parent=self)
def charge(self, **cost):
with self.lock:
if time.monotonic() > self.deadline:
raise BudgetExceeded("deadline")
for k, v in cost.items():
if self.left[k] < v:
raise BudgetExceeded(f"{k} exhausted")
for k, v in cost.items():
self.left[k] -= v
def refund(self, **cost): # settle a reservation that was too high
with self.lock:
for k, v in cost.items():
self.left[k] += v
def note_call(self, tool, args, limit=3):
fp = hashlib.sha256((tool + json.dumps(args, sort_keys=True)).encode()).hexdigest()
with self.lock:
self.seen[fp] = self.seen.get(fp, 0) + 1
if self.seen[fp] > limit:
raise BudgetExceeded(f"loop: {tool} called {self.seen[fp]} times with identical arguments")The gateway step that uses it caps fan-out, reserves and settles, and replaces oversized output with a truncated view plus a handle the model can page through if it really needs more:
MAX_OUTPUT_BYTES = 16_000 # what may enter the model context per call
MAX_FANOUT_PER_STEP = 8 # parallel tool calls the model may request in one step
def run_step(budget, tool_calls, registry, blob_store):
if len(tool_calls) > MAX_FANOUT_PER_STEP:
return [error_result(tc, f"refused: {len(tool_calls)} calls in one step, limit {MAX_FANOUT_PER_STEP}")
for tc in tool_calls]
results = []
for tc in tool_calls:
tool = registry[tc.name]
try:
budget.note_call(tc.name, tc.args)
budget.charge(calls=1, dollars=tool.max_price) # reserve the worst case
raw = tool.invoke(tc.args, timeout=tool.timeout_s) # ONE retry layer lives here or nowhere
budget.refund(dollars=tool.max_price - tool.actual_price(raw))
except BudgetExceeded as e:
results.append(error_result(tc, f"budget: {e}"))
break # stop the step, keep partial results
body = raw.encode() if isinstance(raw, str) else raw
if len(body) > MAX_OUTPUT_BYTES:
handle = blob_store.put(body) # full output kept out of band
body = body[:MAX_OUTPUT_BYTES] + (
f"[truncated {len(body)} bytes; read more with read_blob(handle='{handle}', offset=N)]").encode()
results.append(ok_result(tc, body.decode(errors="replace")))
return resultsThe fingerprint only catches the simplest loops; add a no-progress detector that stops the run when several consecutive steps add no new URLs, files or results. And scope truncation handles to the run and user, or the blob store becomes a cross-tenant leak.
Worked example: the research agent and the link farm
A research agent answers questions by searching, fetching pages and delegating subtopics to sub-agents. Its defaults look reasonable: up to ten parallel fetches per step, up to three sub-agents per task, two levels of delegation and three retries per fetch. A user asks it to summarise opinions about a product. One of the search results is a page that lists 120 review links and a sentence asking any assistant to read every review and to check page two of the list, which links to page three.
Without budgets, the planner fans out ten fetches per step for twelve steps, appending review pages of about 6,000 tokens each. It also delegates three subtopics; each sub-agent finds the same page and repeats the pattern, and delegates again to depth two, giving thirteen agents, all following page two and three. That is close to 2,000 tool calls and, with transcripts of hundreds of thousands of tokens, tens of millions of input tokens.
With the ledger the run is bounded from the start: 60 calls, 400,000 tokens, two dollars and five minutes for the whole tree. The fan-out cap stops the first step at eight fetches. The output cap stores each review page out of band and puts 4,000 tokens into context. The planner keeps 20 calls and carves 25 and 15 for two sub-agents; a third request is refused because nothing is left to carve. The page-two loop trips no fingerprint, because the arguments differ, but the no-progress detector fires after three steps that add only near-duplicate reviews. The agent returns a summary with a note that it stopped early and why. Total spend: 58 calls and well under the token budget, with an answer the user can act on.
Detection and monitoring
Record per run the calls by tool, maximum depth, peak fan-out, output bytes, tokens, dollars and the reason the run ended. Alert on these signals:
- Amplification ratio: tool calls or tokens per kilobyte of input. A jump for one tenant or source domain usually means a planted instruction or a changed tool.
- Budget exhaustion by reason. Rising loop hits point to a broken tool.
- Output bytes per tool over time, which shows an upstream returning full documents before the bill does.
- Retry counts per upstream. Many retries against one host is a retry storm in progress; open a circuit breaker for that tool across all runs.
- Runs killed with side effects still pending, such as started jobs or created resources. These need clean-up, covered by Agent Kill Switch, in depth.
Failure modes and trade-offs
Budgets that are too tight break legitimate work. A data-migration agent may need hundreds of calls; a research agent may genuinely need deep delegation. Set budgets per task class from observed distributions, for example the 99th percentile of successful runs plus headroom, and give users an explicit way to request a larger budget with approval, rather than raising the default for everyone.
Side effects escape budgets. A tool that starts an asynchronous job or creates a VM returns quickly while the real cost continues. Price such tools at their real cost, require approval above a threshold, and tie every resource they create to a run-scoped lease. Running tools inside a sandbox with its own CPU, memory and network quotas, as described in Agent tool-execution sandboxing, bounds what a single call can consume.
Caching identical fetches inside a run removes much of the cost of loops, but a cache shared across users needs keys that include the user's authorisation scope. Truncation hides information, so the marker must say that more exists and how to read it. Budget errors are model input too; word them as facts and never include fetched content in them.
What to do next
- Inventory every tool with its worst-case price, latency and output size, and mark the ones with side effects that outlive a call.
- Put a budget ledger in the tool gateway with calls, tokens, dollars, wall clock, depth and per-step fan-out, enforced outside the model.
- Make sub-agent budgets carved from the parent, and set a delegation depth limit of one or two unless a task class needs more.
- Cap output bytes per call at the boundary, store the rest out of band behind a run-scoped handle, and test with a tool that returns 10 MB.
- Collapse retries to a single layer, charge every retry to the budget, and add a circuit breaker per upstream.
- Add identical-call fingerprinting and a no-progress detector, and return budget exhaustion to the model as a structured result so it can report partial work.
- Alert on the amplification ratio, exhaustion reasons and output bytes per tool, and rehearse a red-team input that asks the agent to fetch every link on a page.