What admission control is, precisely
Model the server as a queueing system. Requests arrive at rate λ (requests/second), the server completes them at rate μ (requests/second), and L requests are in the system — queued plus in service. An admission controller is a predicate evaluated per arrival:
admit(r) = true if predicted_latency(r | current state) ≤ SLO
= false otherwise → reject (429) or deferThe key word is predicted: the decision uses only state observable at arrival time — queue depth, in-flight work, free KV memory — and a model of how long this request will take given that state. It is not scheduling (choosing which admitted request runs next) and not backpressure (telling upstreams to slow down); it is the binary gate in front of both. Everything that follows is about making that predicate cheap, honest, and derived from the SLO rather than from a guessed magic number.
Little's law: the bridge from SLO to queue depth
Little’s law holds for any stable system, with no assumptions about arrival distributions:
L = λ · W
L : average number in system
λ : arrival (= departure) rate, req/s
W : average time in system, secondsIts power for admission control is that it runs backwards. You do not control W directly — latency is an outcome. But you can control L by refusing arrivals. If your SLO says time-in-system must average at most W_max and your service can sustain throughput λ, then the number of requests you may allow inside is bounded:
L_max = λ · W_maxA server that completes 10 req/s with a 2 s SLO should hold at most about L_max = 10 × 2 = 20 requests. The twenty-first admitted request does not get served faster because you were generous — it drags every average up past the SLO. Queue-depth caps are Little’s law wearing a config file.
Why the queue explodes: utilization and 1/(1-rho)
Little’s law bounds the average; the M/M/1 model shows how violently the average moves. Define utilization ρ = λ/μ. For an M/M/1 queue:
W = 1 / (μ − λ) (mean time in system)
L = ρ / (1 − ρ) (mean number in system)Both blow up as ρ → 1. Take μ = 10 req/s, service time 100 ms. At ρ = 0.5, W = 1/(10−5) = 200 ms. At ρ = 0.9, W = 1 s. At ρ = 0.99, W = 10 s — a 50× latency penalty for serving 2× the traffic of the half-loaded case. The curve is a hyperbola, not a line, so “we have 10% headroom” near saturation means almost nothing. Admission control’s job is to pin the operating point on the flat part of the curve, typically ρ ≤ 0.7–0.8, by shedding the arrivals that would push it up the wall.
LLM twist 1: service times are long and wildly variable
Classic web requests take milliseconds with modest variance. An LLM request’s service demand is roughly:
T(r) ≈ T_prefill(n_in) + n_out · t_decode
n_in : prompt tokens (known at arrival)
n_out : output tokens (unknown at arrival!)
t_decode : seconds per generated tokenTwo problems. First, T(r) spans two to three orders of magnitude — a 20-token reply and a 2,000-token reply are both “one request.” Queueing theory says waiting time grows with the variance of service time (the Pollaczek–Khinchine formula has a (1 + C_v^2) factor, where C_v is the coefficient of variation), so LLM queues are intrinsically worse than their mean suggests. Second, n_out is unknown, so the admission predicate must use an estimate — the user’s max_tokens, a per-route historical mean, or a small predictor — and should be conservative, because underestimating admits work you cannot finish in time.
LLM twist 2: the scarce resource is KV memory, not CPU alone
A continuous-batching server admits a request into a running batch only if there is KV-cache room for it. Per-request KV footprint:
KV(r) = 2 · n_layers · n_kv_heads · d_head · (n_in + n_out) · bytes
Example: 32 layers, 8 KV heads, d_head = 128, fp16 (2 B):
KV per token = 2 · 32 · 8 · 128 · 2 ≈ 131 KB
2,048-token request ≈ 268 MBWith, say, 8 GB of KV budget, at most ~30 such requests can be resident regardless of how fast the arithmetic is. Admission control therefore has a second predicate: free_KV ≥ KV_reserved(r), where the reservation uses the worst-case n_in + max_tokens unless the engine supports preemption. Admitting on optimistic memory math causes mid-generation eviction or swap — the most expensive possible failure, because the work already done is thrown away or stalls everyone else.
Deriving the admission threshold from the SLO
Put the pieces together. Suppose the SLO is on time to first token (TTFT), the number users feel most. A new arrival must wait for queued prefill work ahead of it, then its own prefill:
TTFT(r) ≈ Σ_{q in queue} T_prefill(q) / capacity + T_prefill(r)
admit(r) ⇔ TTFT(r) ≤ TTFT_SLO and free_KV ≥ KV_reserved(r)In practice you rarely sum per-request estimates; you track a single scalar — queued tokens — and divide by measured prefill throughput (tokens/s) to get expected queue delay. That gives a token-denominated threshold instead of a request count, which is the right unit when request sizes vary 100×. Note what this is not: it is not a static rate limit (“100 req/min per key”), which protects fairness but knows nothing about current load. Admission control is load-aware by construction; the threshold moves as the queue and memory state move.
A worked example, end to end
Concrete numbers. A CPU SLM server sustains prefill at 1,200 tokens/s and decodes at 40 tokens/s aggregate. SLO: TTFT ≤ 2 s at p50. A new request arrives with an 800-token prompt.
Own prefill: 800 / 1200 = 0.67 s
Budget for queue: 2.0 − 0.67 = 1.33 s
Max queued toks: 1.33 × 1200 ≈ 1,600 tokens
Currently queued: 1,100 tokens → wait ≈ 0.92 s
Predicted TTFT: 0.92 + 0.67 = 1.59 s ≤ 2 s → ADMIT
If queued = 2,400: wait = 2.0 s, TTFT = 2.67 s → REJECTThe controller never measured latency directly — it converted the SLO into a token budget once, then compared one counter against it per arrival, an O(1) decision. Add the memory gate: if the request reserves 800 + 512 = 1,312 tokens of KV at 131 KB/token ≈ 172 MB and only 150 MB is free, it is rejected even though the latency check passed. Both predicates must hold.
Reject fast: why 429 beats a slow timeout
Rejecting feels like failure, so teams let the queue absorb overload instead. The math says that is the crueler choice. Suppose capacity is μ = 10 req/s, offered load is λ = 15 req/s, and clients time out at 10 s. Without admission control the queue grows at 5 req/s; within seconds every position beyond 10 × 10 = 100 deep is doomed to time out. The server then spends real prefill compute on requests whose clients have already hung up, so goodput — completed-within-deadline work — falls below μ even though the machine is 100% busy. That is congestion collapse.
With a cap at L_max = λ_srv · W_max, the server does 10 req/s of useful work forever and returns the excess 5 req/s a 429 in microseconds. Every admitted request meets its SLO; every rejected one learns instantly and can retry elsewhere. A fast no preserves goodput; a slow yes destroys it.
Retry amplification: pricing the rejection itself
Rejections are not free — rejected clients retry. If a fraction p of offered load is rejected and every rejection retries immediately, effective offered load becomes:
λ_eff = λ · (1 + p + p^2 + …) = λ / (1 − p)Rejecting 50% of traffic that retries instantly doubles arrivals — the controller manufactures its own overload. Two standard fixes change the math. Exponential backoff with jitter spreads the geometric series over time so λ_eff stays near λ at any instant. Retry-After turns rejection into deferral: the server computes when capacity should exist — roughly queued_tokens / throughput seconds ahead — and tells the client. A useful client-side companion is the retry budget (e.g. retries ≤ 10% of requests), which caps p’s amplification no matter how the server behaves. Admission control design includes the reject path’s arithmetic, not just the admit path’s.