What the name actually denotes
Strip the branding and in-flight batching is iteration-level scheduling: the serving loop revisits the batch after every forward pass, evicts requests that just emitted their end-of-sequence token, and admits waiting requests into the freed slots — all without draining the batch first. That mechanism is not a TensorRT-LLM invention; it is the Orca contribution, and vLLM exposes the same behaviour under the label continuous batching.
So the honest framing is that in-flight batching and continuous batching name the same scheduling discipline, and much of what is true of one is true of the other. The reason TensorRT-LLM keeps a distinct term is that the discipline is fused into a compiled, kernel-level runtime rather than a Python loop over eager kernels. The interesting content of this article is therefore not the scheduling idea — the sibling articles derive its throughput math — but the engine-level pieces that let a static, ahead-of-time-compiled graph behave like a dynamic, per-step scheduler at all.
The runtime that owns the loop
In TensorRT-LLM the serving loop lives in a C++ runtime component — historically the batch manager (GptManager), and in recent versions surfaced through the higher-level Executor API — that sits between incoming requests and the compiled engine. You enqueue requests; the runtime owns the decision of which of them form the batch on each iteration.
Its loop, each step, does three things: it asks a capacity scheduler which pending or in-progress requests can be afforded this iteration given free KV memory, it assembles those into a single packed input, and it launches one engine execution that advances every one of them by the appropriate number of tokens. Completed sequences are returned and their resources released before the next step. Because the engine itself is a fixed graph compiled for a maximum batch size and sequence length, the runtime’s job is to keep feeding that graph well-shaped, fully-populated batches — the scheduler is the dynamic brain wrapped around a static body.
One forward pass over mixed phases
The mechanism that makes it all pay off is that a single engine execution can carry requests in different phases. A newly admitted request needs its whole prompt processed — the context (prefill) phase, which reads many tokens at once. An in-progress request needs exactly one new token — the generation (decode) phase. TensorRT-LLM packs both kinds into the same batch and runs them together.
This works because the GPT attention plugin is written to handle both regimes. Inputs are stored packed — the runtime removes padding (remove_input_padding) and concatenates every request’s tokens into one long sequence, with an accompanying array of per-request lengths so the kernels know where each begins. A context request contributes many query positions; a decode request contributes one. The attention kernel dispatches the context-phase math for the former and the generation-phase (single-query) math for the latter, reading each request’s history from its own KV blocks. No token budget is wasted on padding, and prompts and steady-state decodes share the pass.
Paged KV and the block manager
Iteration-level scheduling is only useful if you can hand memory to a new request and reclaim it from a finished one cheaply, mid-flight. A contiguous per-request KV buffer cannot: it fragments the moment sequences enter and leave at different times. TensorRT-LLM solves this the same way vLLM does, with a paged (block-based) KV cache managed by a dedicated KV cache manager.
The cache is carved into fixed-size blocks, each holding the keys and values for a set number of token positions. A sequence holds a list of blocks rather than one slab, and grows by acquiring a new block only when its current one fills. When a request finishes, its blocks return to the free pool for immediate reuse by whoever is admitted next. This is what lets the scheduler admit and evict every iteration without compaction: memory is allocated at block granularity, so fragmentation is bounded to at most one partly-filled block per sequence, and the manager can also share identical prefix blocks across requests to reuse cached context.
Capacity scheduler policies
The scheduler must decide how aggressively to fill the batch, and TensorRT-LLM exposes this as a choice of capacity scheduling policy. The two that matter are MAX_UTILIZATION and GUARANTEED_NO_EVICT, and they trade throughput against predictability.
MAX_UTILIZATION packs in as many requests as the KV memory can hold right now, betting that most will finish before the cache is exhausted. It maximizes concurrency and tokens per second, but it can over-commit: if too many admitted sequences keep growing, the runtime must pause and evict one, saving or recomputing its state later. GUARANTEED_NO_EVICT is the conservative policy — it only admits a request if there is enough KV memory to see it through to its maximum length, so once a request starts it is never preempted. That costs some peak throughput (the batch runs a little emptier) but removes eviction stalls and the tail-latency jitter they cause. The right choice depends on whether you are optimizing aggregate throughput or per-request latency guarantees.
Chunked context, briefly
A long prompt is a problem for a mixed batch: its context phase can be so large that the iteration it lands in becomes far heavier than a normal decode step, stalling every request sharing that pass. TensorRT-LLM addresses this with chunked context — splitting a prompt’s prefill across several iterations so each contributes only a bounded slice of tokens.
The dedicated sibling article covers the mechanics; what matters here is how it composes with in-flight batching. Chunking turns an otherwise spiky context request into a stream of uniform-sized pieces that the iteration scheduler can interleave with ongoing decodes, keeping every forward pass roughly the same shape. In-flight batching supplies the per-iteration admission machinery; chunked context supplies the token-budget discipline that keeps those iterations balanced. Together they prevent a single 8k-token prompt from freezing the decode stream of everyone else in the batch — the head-of-line blocking that naive prefill-first batching suffers.
A scheduling walkthrough
To see the policies behave, picture an engine built for a max batch of 8 sequences with a KV pool of 100 blocks, each block holding 16 token positions. Six requests are decoding, together holding 70 blocks. Two new requests arrive, one a short 200-token prompt (needs ~13 blocks now, growing), one a 1500-token prompt (needs ~94 blocks at full length).
Under GUARANTEED_NO_EVICT, the runtime admits the short prompt — 13 blocks fit in the free 30 with headroom for its growth — but holds the long one, because it cannot reserve the ~94 blocks that request could eventually demand. Under MAX_UTILIZATION, it may admit the long prompt too, processing its context and betting some of the six decoders finish before the pool runs dry; if that bet fails, it pauses the lowest-priority sequence and frees its blocks. Same batch, same instant — two policies, two different concurrency levels and two different tail-latency profiles.