An agent run is a job, not a request
Deploying an agent with the same architecture you'd deploy a normal API service -- a stateless container behind a load balancer, autoscaled on request rate, request killed and retried on timeout -- breaks on contact with what an agent run actually looks like: it can take seconds or it can take twenty minutes, it accumulates state across many tool calls that shouldn't be thrown away on a transient failure, and if it misbehaves it can take an action with real consequences, not just return a wrong HTTP response. This article works through the deployment-architecture decisions that follow from taking those differences seriously.
Sandboxing execution
An agent that can execute code, run shell commands, or make arbitrary tool calls needs an execution boundary between what it can touch and what it can't -- the same principle as permission boundaries, applied at the infrastructure layer rather than the credential layer.
Container-per-run gives each agent run its own fresh, isolated container: strongest isolation (a compromised or buggy run can't affect another run, because there's no shared state to affect), but container startup latency (typically hundreds of milliseconds to a few seconds, depending on image size and cold-start handling) is pure overhead on every run, and idle capacity sits unused between runs unless the orchestration layer is tuned to scale containers down aggressively.
A shared sandbox pool (a warm pool of pre-started, reusable execution environments, reset between runs) amortizes the startup cost across many runs, at the price of needing to actually trust the reset step -- any state that leaks across a reset (a lingering process, a modified filesystem, a cached credential) is a cross-run contamination bug, and unlike container-per-run, that class of bug is silent until something specifically goes looking for it. Pools are the right choice when run volume is high enough that per-run container startup becomes a real cost or latency line item; container-per-run is the right default otherwise, because its failure mode (slightly higher latency) is far less dangerous than the pool's (silent state leakage).
Why agent runs break request-response autoscaling
A normal service autoscales on a proxy for load (requests per second, CPU) because request duration is roughly uniform and short -- ten thousand requests per second at 50ms each is a predictable, steady amount of concurrent work. An agent run's duration varies by two orders of magnitude depending on how many steps the task actually needs, so "requests per second" stops being a useful proxy for how much concurrent capacity is actually in use; ten agent runs started in the same second could mean ten seconds of total work or ten minutes, and the autoscaler needs to react to concurrent in-flight runs and their actual resource consumption, not arrival rate.
The practical fix is to treat agent execution as a job queue rather than a request-response endpoint: an API call enqueues a run and returns a handle immediately (or streams progress), a worker pool pulls from the queue and executes, and the worker pool autoscales on queue depth and worker utilization -- the same shape as a batch-processing system, not a web service, because that's what an agent run's duration profile actually resembles.