Most teams discover LLM cost the same way: a monthly invoice that is larger than the spreadsheet said it would be. The usual reaction is to switch to a cheaper model, which sometimes works and often makes things worse, because the cheaper model fails more tasks and every failure is paid for twice. Cost optimization done properly is not a model choice. It is an accounting exercise followed by a short list of prompt-level changes applied in a deliberate order.
This article builds that accounting from first principles: what a single request costs, why the number that matters is cost per successful task, and which levers move which terms. It then works one realistic example through with arithmetic, from about $1,955 a day to about $325, and closes with guardrails and a checklist. Prices quoted are Anthropic list prices checked on 2026-10-02; they change, so treat them as an example of the method, not as constants.
The cost equation for one request
Every hosted model bills by tokens, split into categories that are priced differently. For a request against a provider with prompt caching, the cost is the sum of four terms: uncached input tokens at the input price, tokens written to the cache at a premium, tokens read from the cache at a discount, and output tokens at the output price. On Anthropic's API the cache write premium is 1.25 times the input price for the default five-minute lifetime and 2 times for the one-hour lifetime, and a cache read is about 0.1 times the input price on most models. Output is typically several times the input price per token; for Claude Sonnet 5.5 the list prices are $2 per million input tokens and $10 per million output tokens.
Two consequences follow directly. First, an output token is worth five input tokens, so a prompt that makes the model write half as much often saves more than a prompt that is half as long. Second, reasoning or thinking tokens are billed as output even when you never display them, so a route that thinks hard about an easy question carries an invisible tax.
Measure before you cut
You cannot optimize what you have not attributed. Every response carries a usage object, and on Anthropic's API it reports input_tokens, cache_creation_input_tokens, cache_read_input_tokens and output_tokens. Turn those into dollars at the call site and log them with the route, the prompt version, a task identifier and the attempt number. That is enough to answer the three questions every cost review asks: which route spends the money, which prompt version changed it, and how much is spent on attempts that did not succeed.
from dataclasses import dataclass
@dataclass
class Price: # USD per million tokens; date-stamp these, they change
inp: float
out: float
write_mult: float = 1.25 # 5-minute cache write (1-hour TTL: 2.0)
read_mult: float = 0.10 # cache read; check per model
SONNET_5_5 = Price(inp=2.00, out=10.00) # list price checked 2026-10-02
def request_cost(usage, price: Price) -> float:
"""usage is the response.usage object (or a dict of the same fields)."""
g = usage.get if isinstance(usage, dict) else lambda k, d=0: getattr(usage, k, d) or d
m = 1_000_000
return (
g("input_tokens", 0) * price.inp / m
+ g("cache_creation_input_tokens", 0) * price.inp * price.write_mult / m
+ g("cache_read_input_tokens", 0) * price.inp * price.read_mult / m
+ g("output_tokens", 0) * price.out / m # thinking tokens bill as output
)
def log_call(route, prompt_version, task_id, attempt, usage, price, ok):
record = {
"route": route, "prompt_version": prompt_version, "task_id": task_id,
"attempt": attempt, "cost_usd": round(request_cost(usage, price), 6),
"in": usage.input_tokens, "cache_read": usage.cache_read_input_tokens or 0,
"cache_write": usage.cache_creation_input_tokens or 0,
"out": usage.output_tokens, "ok": ok,
}
emit_metric(record) # your log pipelineKeep the price table in one place with a date on it. When prices change, the history should stay priced at what you actually paid, so store the computed cost per call rather than recomputing it later from token counts.
Cost per successful task, not cost per request
A request is not the unit of value. A task is: a ticket triaged, an answer accepted, an extraction that passed validation. A task may take several requests: a retry after malformed JSON, a second turn because the first answer was incomplete, an agent loop with eight tool calls. The honest metric is total spend divided by tasks that passed your quality check.
This is the number that stops bad optimizations. Shortening a prompt so much that the failure rate climbs from 2% to 15% lowers cost per request and raises cost per success. Moving a hard route to a smaller model can triple the turns an agent needs. Compute the metric on a fixed evaluation set so that two prompt versions are compared on the same inputs, as described in prompt evaluation architecture.
import collections
def cost_per_success(records):
"""records: one dict per API call, grouped by task_id; ok marks the final verdict."""
spend = collections.defaultdict(float)
passed = {}
for r in records:
spend[r["task_id"]] += r["cost_usd"]
if r["ok"]:
passed[r["task_id"]] = True
total = sum(spend.values())
wins = len(passed)
return total / wins if wins else float("inf")
# Compare two prompt versions on the SAME eval set, never on live traffic mixes
a = cost_per_success(run_eval(prompt="triage_v7"))
b = cost_per_success(run_eval(prompt="triage_v8_short_output"))
Lever 1: cache the stable prefix
Most production prompts are mostly constant: a role and policy block, label definitions, tool schemas and a fixed set of examples, followed by a small variable part. Caching charges full price for that constant part once per cache lifetime and about a tenth of the input price afterwards. With the five-minute lifetime, two requests already break even: 1.25 plus 0.1 is less than 2.
Caching is a prefix match, so the work is structural. Put frozen content first and volatile content last, keep tool definitions in a deterministic order, and never let a timestamp, request ID or per-user detail leak into the system block. Anthropic allows up to four breakpoints per request and has a model-dependent minimum cacheable length, between 512 and 4,096 tokens, below which nothing is cached and no error is raised. The mechanics are covered in prompt caching architecture; the cost discipline is simply to check that cache_read_input_tokens is non-zero on the second call of every route.
# Order the request so the stable bytes come first and never change between calls
system = [
{"type": "text", "text": POLICY_AND_ROLE}, # frozen
{"type": "text", "text": LABEL_DEFINITIONS}, # frozen
{"type": "text", "text": FEW_SHOT_EXAMPLES, # frozen, fixed order
"cache_control": {"type": "ephemeral"}}, # breakpoint after the static part
]
messages = [{"role": "user", "content": render_ticket(ticket)}] # volatile, after the breakpoint
resp = client.messages.create(model="claude-sonnet-5-5", max_tokens=400,
system=system, tools=TOOLS_SORTED, messages=messages)
assert (resp.usage.cache_read_input_tokens or 0) > 0 or is_first_call_in_window()
Lever 2: input hygiene
After caching, the input you still pay full price for is the dynamic part: retrieved passages, conversation history, tool results. Three habits remove most of the waste. Retrieve fewer, better passages: going from top-10 to top-4 with a reranker usually keeps answer quality and removes thousands of tokens per call. Bound history: keep the last few turns verbatim and replace older turns with a short running summary. Trim tool results before they re-enter the context, because a 20 KB JSON response that the model needs three fields from is paid for on every later turn of the loop.
For techniques that compress text without losing the facts the model needs, see prompt compression architecture. Measure each cut with the eval set; context removal fails quietly, as an answer that is fluent but wrong.
Lever 3: an output contract
Because output is the expensive side, the prompt should say exactly what to produce and nothing else. Replace open instructions such as 'analyse this ticket' with a contract: the fields, their allowed values, a maximum length for free-text fields and an explicit instruction not to restate the input. Set max_tokens to a ceiling a little above the longest legitimate answer so that runaway generations are cut off.
Structured output earns its place here twice. A schema-constrained response is shorter than prose that describes the same decision, and it removes the parse-failure retries that silently double the cost of some routes. Structured output with JSON Schema covers the contract itself. Where the model supports a reasoning effort control, set it per route: classification and extraction usually hold quality at low effort, while multi-step reasoning may need more.
Lever 4: batch what can wait
Anthropic's Message Batches API runs requests asynchronously at 50% of standard prices, with results that may arrive in any order and must be matched by a custom ID. Nightly re-scoring, back-fills, evaluation runs, document enrichment and any queue where minutes of delay are invisible to a user belong there. Batch is a trade of latency for money, so route on a field in the job, not on a guess: urgent tickets go synchronous, the rest go to the next batch.
Lever 5: model choice and routing, last
Routing easy requests to a smaller model is the lever people reach for first and should reach for last. It trades quality for money, it needs a classifier that is itself a source of errors, and it splits your cache because cached prefixes are per model. Before building a cascade, measure the simpler alternative: the stronger model with a tighter output contract and a lower effort setting. If routing still wins on cost per success, design it as described in prompt routing architecture, with a fallback path for the small model's low-confidence answers.
Worked example: support-ticket triage
A support platform triages 100,000 tickets a day with Claude Sonnet 5.5 at $2 and $10 per million tokens. The prompt has 6,000 tokens of policy, label definitions and examples, plus an 800-token ticket. The model writes a 450-token response, including a paragraph that restates the ticket, and 8% of responses fail JSON parsing and are retried.
| Step | Per request | Attempts per success | Per day |
|---|---|---|---|
| Baseline: 6,800 input at $2/M, 450 output at $10/M | $0.0136 + $0.0045 = $0.0181 | 1.08 | about $1,955 |
| Cache the 6,000-token prefix (reads at $0.20/M) | $0.0012 + $0.0016 + $0.0045 = $0.0073 | 1.08 | about $788 |
| Output contract: 180 tokens, no restatement | $0.0028 + $0.0018 = $0.0046 | 1.08 | about $497 |
| Schema-constrained output: retries fall to 1% | $0.0046 | 1.01 | about $465 |
| Batch the 60% that are not urgent (50% off) | blended 0.7x | 1.01 | about $325 |
Cache writes are left out of the arithmetic because at this volume a request arrives every second and the prefix is written roughly once per five minutes. The batch row assumes cache behaviour holds inside batches, which you should measure rather than assume. The point of the table is the order: the first three rows are free wins that also improve reliability, and together they remove about three quarters of the bill before any quality trade-off is made. Each row was accepted only after the triage eval showed accuracy within its tolerance.
Guardrails that keep it cheap
- A per-route budget alert on cost per successful task, not on total spend, so a quality regression shows up as a cost signal.
- A cost regression test in CI: run the eval set for each prompt change and fail the build if cost per success rises more than an agreed percentage.
- A cache-hit check per route; a sudden drop to zero cache reads almost always means someone added a timestamp or reordered tools.
- Hard caps on agent loops: a maximum number of turns and a token budget per task, with a clean failure when either is reached.
- Per-tenant quotas so one customer's runaway integration cannot spend the month's budget in an afternoon.
Failure modes and trade-offs
- Optimizing per request. A shorter prompt with more failures costs more per task. Always report cost per success.
- Silent cache misses. A dynamic value in the prefix turns every call into a cache write at 1.25 times the price; the bill goes up after 'adding caching'.
- Truncated answers. A
max_tokensset too low cuts valid answers mid-field, which then fail validation and retry. - Batch for interactive work. Users notice delays measured in minutes; batch only queues whose consumers are machines or next-day reports.
- Routing errors. A misrouted hard request costs a failed small-model attempt plus the large-model retry. Measure the misroute rate before you count savings.
The general trade-off is that caching, input hygiene and output contracts mostly cost engineering time, while batch and routing cost latency and quality risk. Spend the engineering time first.