Every model has a context window: the maximum number of tokens it can attend to in one request. A single prompt rarely hits it. Conversations and agent loops do, because each turn adds the user's message, the model's reply, tool calls and tool results, and nothing ever leaves on its own. Sooner or later a request is rejected for being too long, or, worse, it is silently cut by a framework, or the model is technically within the limit but has no room left to answer.

Context window management is the run-time controller that prevents all three. It decides, before every request, what stays verbatim, what gets condensed, what is dropped and how much room is left for the answer. This page builds that controller from first principles: what counts against the window, how to count it exactly, how to set a budget, an ordered ladder of overflow steps with working code, how to compact without destroying prompt caching, how agents differ, a worked example with real numbers, and the failure modes that show up in production. How to lay out documents inside one long prompt is a separate topic, covered in long context prompting.

Advertisement

What actually counts against the window

Providers and runtimes count input and output against the same window. If a model has a 32,000-token window and you request up to 4,000 output tokens, your input must fit in about 28,000. Some APIs reject a request whose input plus requested maximum output exceeds the window; others accept it and stop generation early. Either way, an answer that needs 4,000 tokens cannot be produced from a 31,000-token input.

The input is more than the visible chat: the system prompt, tool schemas (often thousands of tokens for an agent), the chat template's role markers, every previous message, tool calls and results, retrieved documents, images or audio as tokens, and any thinking text you send back.

Treat these as blocks with different lifetimes. The system prompt and tool schemas are fixed for the session. Recent turns are needed verbatim, older turns only for their facts. Tool results are usually needed once, and retrieved chunks only for the current question.

One request, one window: every token competes for the same budgetsystem promptfixed, cachedtool schemasfixed, cachedsummaryof old turnsrecent turnsverbatimoutput reservemax tokensstable prefixgrows every turnmust stay freecount tokensmodel tokenizer + templateover budget?input + reserve + marginsend requestlog tokens usednoyesOverflow ladder: cheapest, least lossy step first; recount after each step1 clear oldtool results2 drop low-scoreretrieved chunks3 summariseoldest turns4 truncate orstop and askFixed blocks never move, so the provider's prompt cache keeps matching them.
The window is shared by a fixed prefix, a growing middle and an output reserve. Before each request the controller counts, and if the input is over budget it walks an ordered ladder of overflow steps, recounting after each, from cheap and nearly lossless to lossy.

Count tokens the way the model counts them

Every decision below depends on an accurate count, and the only accurate count comes from the model's own tokenizer applied to the fully rendered prompt, template included. Different model families tokenize the same text differently, so a tokenizer for one family can be badly off for another, especially for code, non-English text and numbers. A rule of thumb such as four characters per token is fine for capacity planning and wrong for a guard that must never overflow.

from transformers import AutoTokenizer

tok = AutoTokenizer.from_pretrained("your-org/your-model")

def count_messages(messages, tools=None):
    """Exact input tokens as the model will see them, including role markers."""
    enc = tok.apply_chat_template(messages, tools=tools, add_generation_prompt=True,
                                  tokenize=True, return_dict=True)
    return len(enc["input_ids"])

# For a hosted API, call the provider's token-counting endpoint if it has one,
# or read the usage block of each response and calibrate a local estimate:
def estimate(text, chars_per_token=3.5):
    return int(len(text) / chars_per_token) + 1   # an estimate, never a guarantee

For hosted models, use the provider's token-counting facility if it offers one, and always log the input token count each response reports; comparing it with your estimate tells you how big the safety margin must be.

Advertisement

Set the budget before you need it

Write the budget down as arithmetic, not as a feeling. The input limit is the window minus the output reserve minus a safety margin. The output reserve is the largest answer you will accept, which is also the maximum output tokens you send with the request. The margin absorbs estimate error and template changes; 2 to 5 percent of the window is a reasonable start, tightened once your logs show how accurate the counts are.

Then split the input limit into soft targets per block: fixed prefix, summary cap, retrieval cap, and the rest for verbatim turns. Protect the last few exchanges unconditionally and make everything older negotiable.

You need not use the whole window. Cost and latency grow with input length, and many models use distant content less reliably, so a working budget of 50 to 60 percent of the window often gives better answers and lower bills.

The overflow ladder

When the rendered input exceeds the limit, apply reductions in order from cheapest and least lossy to most lossy, recounting after each one and stopping as soon as the request fits. The order matters more than any individual technique.

  1. Clear old tool results. A result already acted on is the largest, least valuable block. Replace it with a placeholder, keeping the call so the model knows it can call again.
  2. Drop low-scoring retrieval. Drop the weakest chunks first.
  3. Summarise the oldest turns. Fold a block of old turns into a rolling summary. This is lossy, so it comes after the nearly free steps.
  4. Truncate, then stop. Remove the oldest turns, never splitting a tool call from its result. If the protected turns still do not fit, ask the user to narrow the request.
from dataclasses import dataclass, field

@dataclass
class Budget:
    window: int            # model context length
    reserve_output: int    # max_tokens you will request
    margin: int = 512      # tokenizer drift, template changes

    @property
    def input_limit(self):
        return self.window - self.reserve_output - self.margin

@dataclass
class Context:
    system: list           # fixed messages: instructions, tool schemas
    summary: str = ""      # rolling summary of compacted turns
    turns: list = field(default_factory=list)       # verbatim history
    retrieved: list = field(default_factory=list)   # (score, text) chunks

class ContextOverflow(Exception):
    pass

def render(ctx):
    """Stable parts first; per-question retrieval goes last, inside the latest user turn."""
    msgs = list(ctx.system)
    if ctx.summary:      # changes only when compaction runs
        msgs.append({"role": "user", "content": "Summary of earlier conversation:\n" + ctx.summary})
        msgs.append({"role": "assistant", "content": "Understood."})
    turns = [dict(m) for m in ctx.turns]
    if ctx.retrieved and turns and turns[-1]["role"] == "user":
        docs = "".join("<doc>" + chunk + "</doc>" for _, chunk in ctx.retrieved)
        turns[-1]["content"] = docs + "\n" + turns[-1]["content"]
    return msgs + turns

def fit(ctx, budget, count, summarise, keep_recent=6):
    """Apply the cheapest step that still fails the budget; recount after each."""
    steps = [clear_tool_results, drop_retrieval, compact_oldest, truncate_oldest]
    for step in steps:
        while count(render(ctx)) > budget.input_limit:
            if not step(ctx, summarise, keep_recent):
                break                    # this step has nothing left; try the next
        if count(render(ctx)) <= budget.input_limit:
            return ctx
    raise ContextOverflow("cannot fit even the protected turns; ask the user to narrow")

def clear_tool_results(ctx, _s, keep_recent):
    old = ctx.turns[:-keep_recent]
    for m in old:
        if m["role"] == "tool" and not m.get("cleared"):
            m["content"], m["cleared"] = "[tool result removed; call the tool again if needed]", True
            return True
    return False

def drop_retrieval(ctx, _s, _k):
    if not ctx.retrieved:
        return False
    ctx.retrieved.sort(key=lambda sc: sc[0])
    ctx.retrieved.pop(0)                 # lowest score first
    return True

def compact_oldest(ctx, summarise, keep_recent, chunk=8):
    old = ctx.turns[:-keep_recent]
    if len(old) < 2:
        return False
    n = min(chunk, len(old))
    while n < len(ctx.turns) - 1 and ctx.turns[n]["role"] == "tool":
        n += 1                           # keep each tool call with its result
    ctx.summary = summarise(ctx.summary, ctx.turns[:n])   # absorbs a whole block
    del ctx.turns[:n]
    return True

def truncate_oldest(ctx, _s, keep_recent):
    if len(ctx.turns) <= keep_recent:
        return False
    del ctx.turns[0]
    while ctx.turns and ctx.turns[0]["role"] == "tool":
        del ctx.turns[0]                 # never leave an orphaned tool result
    return True

Three details prevent common bugs. Compaction and truncation never separate a tool result from its call, which most APIs reject. Retrieval is rendered into the latest user turn, after the stable history. And each step reports whether it changed anything, so the loop cannot spin.

Summaries that keep what matters

A generic request to summarise produces a friendly paragraph that loses the order number and the constraint from turn three. Give the summariser an explicit keep-list and a hard length cap, and merge into the existing summary rather than summarising a summary of a summary from scratch.

SUMMARY_PROMPT = """You maintain the running memory of a conversation.
Merge the existing summary with the new messages. Keep, verbatim where possible:
- decisions made and their reasons
- facts the user stated about themselves, their data or constraints
- open questions and promised follow-ups
- identifiers: order numbers, file paths, ticket ids, numbers with units
Drop greetings, restated questions and tool output that was already used.
Write at most 400 words. Use bullet points."""

Evaluate it directly: compact long real transcripts, then ask questions whose answers were only in the compacted part, and compare with the full transcript. Keep the uncompacted transcript in a database or file so a later turn can retrieve an exact detail the summary dropped.

Compact in steps to keep the cache warm

Many providers and serving engines cache the processed form of a prompt prefix, so a request whose beginning matches an earlier request is cheaper and faster. Caching matches on an exact prefix. That has a direct consequence for context management: any change near the start of the prompt invalidates everything after it.

Three habits follow. Keep the fixed blocks first and byte-identical across turns, with no timestamps in the system prompt. Append to history rather than editing it. And compact in large steps at a high-water mark: when input passes about 80 percent of the limit, fold a big block into the summary at once. Dropping one turn per request changes the prefix every request and never hits the cache; stepwise compaction breaks it once. How prefix caching works in detail is covered in prompt caching.

Agents fill the window differently

Agent loops stress the window harder than chat, because one user request can trigger dozens of tool calls. Most growth comes from tool output, so the first control is at the source: make tools return what the model needs, not everything they have. Paginate, return the top rows with a count, return file excerpts with line numbers, and put large artefacts in a file or store and return a handle.

Reasoning models add another block: thinking text. Follow the model's own rules for it. Gemma 4's prompt-formatting guide, for example, says to strip the model's thoughts from previous turns in normal multi-turn conversation, but not to strip them between function calls within a single turn. Applying a blanket rule either way either wastes the window or breaks the model's tool-use loop.

Finally, give long-running agents a notes file or task list outside the window, so compaction loses less. Agentic prompting covers how to instruct agents to keep such notes.

Worked example: a support assistant in a 32K window

A support assistant runs on a model with a 32,768-token window. Answers rarely exceed 800 tokens, so the reserve is 1,500. The margin is 768. The input limit is therefore 30,500. The fixed prefix, system prompt plus six tool schemas, measures 3,900 tokens with the model's tokenizer. Retrieval is capped at four chunks of about 600 tokens, 2,400 in all, and the summary is capped at 600. That leaves about 23,600 tokens for verbatim turns, and the last six messages are protected.

A typical session goes like this. By turn eleven the input is 24,900 tokens, past the 80 percent high-water mark of 24,400. The largest item is an order-history tool result of 6,200 tokens from turn four, already used. Step one clears it and two older results, bringing the input to 15,300: done, no summary needed. At turn twenty-three the input passes 24,400 again with no old tool results left. Step two drops nothing, because the current question needed retrieval. Step three folds turns one to fourteen into a 450-token summary, bringing the input down to 13,800. The prompt cache misses once on turn twenty-three and hits again from turn twenty-four.

Failure modes

SymptomCauseFix
Request rejected as too longCounted with the wrong tokenizer or without the templateCount the rendered prompt with the model's tokenizer; keep a margin
Answer stops mid-sentenceInput left too little room for outputReserve output explicitly and subtract it from the input limit
Model forgets an early constraintTruncation or a vague summary dropped itKeep-list in the summariser; pin hard constraints in the system prompt
API error about a tool resultTruncation split a call from its resultRemove calls and results as pairs
Costs jump after compaction is addedPer-turn trimming changes the prefix every requestCompact in steps at a high-water mark; keep fixed blocks first
Framework silently cuts the startLibrary default truncationOwn the policy; log every compaction

Trade-offs

Verbatim history is exact and expensive; summaries are cheap and lossy, and the loss stays invisible until someone asks about the lost detail. A larger window postpones every decision but raises cost and latency on every turn. For squeezing more into a fixed budget, see prompt compression, and for choosing what to include in the first place, context packing.

What to do next

  1. Log input and output tokens per request, split by block: fixed prefix, summary, retrieval, turns, tool results.
  2. Replace any character-based estimate in your guard with a count of the rendered prompt from the model's tokenizer or the provider's counter.
  3. Write the budget as arithmetic: window, output reserve, margin, and a soft cap for each block.
  4. Implement the overflow ladder in order and test it on your longest real transcripts.
  5. Give the summariser an explicit keep-list and evaluate it with questions about compacted turns.
  6. Make fixed blocks byte-identical across turns and compact at a high-water mark, then confirm cache hits in your usage logs.
  7. Trim tool output at the source and give long-running agents a notes file outside the window.
Key takeaway: The context window is a shared budget for the fixed prefix, the growing history and the answer. Count the rendered prompt with the model's own tokenizer, reserve output explicitly, protect the most recent turns, and when the input is over budget apply an ordered ladder: clear old tool results, drop weak retrieval, summarise old turns with a keep-list, and only then truncate or ask the user to narrow. Compact in large steps with fixed blocks first so prompt caching keeps working, and keep the full transcript outside the window for exact recall.