A gateway is infrastructure; a router is application logic

Once an organization has more than one team calling LLM APIs, a pattern repeats: each team hardcodes its own API keys, its own retry logic, its own fallback-to-a-different-model logic, and its own cost tracking -- because there was no shared layer to put it in. An LLM gateway is that shared layer: a single internal service every application calls instead of calling providers directly, responsible for the concerns that are genuinely the same across every caller (credentials, rate limits, provider outages, cost visibility) so that no individual application has to re-solve them.

This is a different concern than agent request routing, which decides which model or agent should handle a given task based on the task's difficulty -- that's application-layer judgment specific to one product. A gateway sits underneath all of that: every app's router, including the difficulty-based one, still calls out through the gateway to actually reach a model. Confusing the two layers is the most common design mistake in this space -- teams build gateway-shaped logic (fallback, retries, rate limiting) separately inside each application's router instead of once, underneath all of them.

Advertisement

What a gateway is actually responsible for

Credential and rate-limit pooling. Individual teams calling a provider directly each hit their own per-key rate limit, even though the organization's aggregate usage is well under what the provider would grant a single, larger pooled key. A gateway holds the provider credentials centrally and pools quota across all internal callers, so one team's traffic spike doesn't need its own limit-increase request and another team's quiet period doesn't go to waste.

Request routing by cost, latency, and capability. The gateway is where "use the cheap model for this, the frontier model for that" gets enforced as policy rather than left to each application's discretion -- a routing table keyed on a request tag (which an application sets: tier=cheap, tier=quality, or a specific capability requirement like function calling or a long context window) maps to a concrete provider and model, so the mapping can be changed centrally (a cheaper model gets promoted to the default cheap tier) without redeploying every application that uses it.

Provider-outage fallback. When a provider has a degraded period, individual applications calling it directly all fail simultaneously with no coordinated response. A gateway can detect a rising error rate or latency spike from one provider and automatically fail over new requests to a configured secondary provider/model, transparently to the calling application -- the tier abstraction from routing is what makes this possible: "give me a quality-tier model" can be satisfied by more than one provider.

Advertisement

Cost attribution and budgets

Without a gateway, "how much did team X spend on LLM calls last month" requires reconciling provider billing dashboards against application logs after the fact, if it's answerable at all. A gateway sees every request, so it can tag and aggregate cost per caller (team, application, even per end-user if the calling application passes an identifier through) in real time rather than at month-end reconciliation.

This enables the thing cost visibility alone doesn't: enforceable budgets. A per-team monthly token budget, enforced at the gateway, degrades gracefully -- route to a cheaper model once 80% of budget is consumed, hard-stop or require an override once 100% is hit -- rather than the alternative of an unexpected bill arriving after the spending already happened.

request -> gateway
             |-- resolve caller identity + budget status
             |-- resolve tier -> provider/model (routing table)
             |-- check provider health -> fallback if degraded
             |-- check response cache -> return cached if hit
             |-- forward to provider, record cost + latency
             |-- return response, update budget counters