Budget before you build, not after you measure
The common failure pattern: a team builds an LLM product, ships it, and only then discovers it costs four times the acceptable per-request budget or takes six seconds when the product needed sub-two. Both numbers were knowable at design time -- they're the sum of components whose individual costs and latencies are estimable before a line of code exists. Treating cost and latency as a budget to decompose and allocate upfront, the same discipline applied to any other system with a hard SLA, catches the mismatch while it's still a design change instead of a post-launch scramble.
Decomposing end-to-end latency into a budget
Start from the target -- say, a product requirement that a response starts streaming within 2 seconds -- and allocate that total across every component the request actually passes through before the model produces its first token: retrieval (if the request involves RAG, see RAG pipeline design for the retrieve/re-rank latency breakdown), any tool calls that must complete before generation can start, the gateway hop (see LLM gateway architecture), and the model's own time-to-first-token, which itself varies by model size and current provider load.
A concrete allocation for a 2-second budget: 100ms network/gateway overhead, 300ms retrieval, 100ms re-ranking, 1200ms model time-to-first-token, leaving roughly 300ms of margin. The purpose of writing the allocation down explicitly, per component, before building is that it turns "the product feels slow" from a vague post-launch complaint into a specific, attributable question: which component blew its allocated budget, and by how much -- the same diagnostic value a per-span latency budget gives you once you have tracing in place to actually measure each component against its allocation.
Decomposing cost-per-request into a token budget
The same exercise, in tokens and dollars rather than milliseconds. For a RAG-backed request: input tokens are dominated by retrieved context (say, 5 chunks at 400 tokens each, 2000 tokens) plus the system prompt and conversation history (call it 800 tokens), and output tokens are whatever the response actually needs (a support answer might target 150-300 tokens). At current frontier-model pricing in the few-dollars-per-million-tokens range for input and a few times that for output, a single request like this lands in the range of a fraction of a cent to a few cents, and the point of writing out the token breakdown is that it shows exactly which lever moves the number most: for most RAG-backed requests, retrieved context dominates input tokens by a wide margin, which means chunk count and chunk size (the same knobs from RAG pipeline design) are the primary cost lever, not prompt engineering elsewhere in the request.
Multiply the per-request cost by expected volume before committing to an architecture, not after -- a request that costs $0.02 looks trivial in isolation and becomes a real budget line at a million requests a day, and the point where that math needs to happen is during design, when switching to a cheaper retrieval strategy or a smaller model is a design decision, not a post-launch emergency migration.