Why agent-generated code cannot run in your process

The tempting first implementation is four lines long. The model returns a code block, you strip the fences, you call exec(), you feed stdout back into the conversation. It works immediately, demos beautifully, and is indefensible — because exec() inherits everything: your environment variables, which is where your API keys are; your already-authenticated database pool; your cloud SDK’s ambient credentials; and your network namespace, which reaches every internal service your VPC allows.

Python has no meaningful in-language sandbox, and this is not a gap someone will eventually fill. Restricting builtins, blocking import, or scanning the source for dangerous substrings are all defeated by well-known one-liners that walk the object graph from any innocuous object back to the interpreter internals — and AST allowlists fall the same way once you permit the attribute access and function calls any real analysis snippet needs.

ADK marks this honestly: the local executor is named UnsafeLocalCodeExecutor, and the name is the documentation. Treat its presence in a deployment manifest the way you would treat verify=False on a TLS call — a lint failure, not a configuration choice.

Advertisement

The threat model: prompt injection is the exploit primitive

Classic application security assumes the attacker must find a flaw. Here they do not: the flaw is the feature. An LLM cannot reliably distinguish instructions from data, because both arrive as tokens in one undifferentiated context. The attack is not an injection into a parser — it is an injection into the decision maker, and it is written in English.

The dangerous configuration is three properties sharing one trust zone: the agent processes attacker-influenced content, it holds a capability that does real work, and it has a path to send data outward. Any two are survivable; all three is an exfiltration pipeline that runs on demand. And note that tool output is attacker-controlled content — a scraped page, an issue-tracker comment, a row deep inside an uploaded CSV. Your own tool delivered it, which makes it feel trusted; it is not.

The design consequence is worth stating plainly: assume the model will be persuaded. Do not architect for a 99% defense rate; at a thousand sessions a day that is ten compromises. Architect so that a fully hijacked agent emitting the worst code an attacker can write accomplishes nothing, because everything it needs — credentials, reachable hosts, persistent storage — sits on the far side of a boundary it does not control.

Advertisement

The isolation ladder — from exec() to a rented boundary

There is no single right answer, only a ladder trading isolation strength against cost and startup latency. Pick the rung your threat model demands, not the one your convenience prefers.

RungIsolation strengthStartupUse when
In-process exec()None0 msNever in production; local dev only
Subprocess + rlimitsWeak; same kernel, same filesystem, same network~10 msFully trusted input, which you do not have
Container (non-root, dropped caps, read-only root, seccomp)Moderate; shared kernel is the attack surface0.1–1 sThe sensible default for most agents
gVisor / Firecracker microVMStrong; syscalls intercepted in userspace or a separate kernel0.2–2 sMulti-tenant, or genuinely hostile input
Hosted code-execution serviceStrong; someone else patches itProvider-dependentYou would rather rent the boundary than operate it

The honest framing of the container rung: a container is a bundle of kernel features, not a security perimeter, and escapes via kernel bugs are a regular occurrence. That is acceptable single-tenant, where the blast radius is one customer’s own data on a disposable node. It is not acceptable when one node runs code from many customers — which is exactly what gVisor and microVMs exist for, and a few hundred milliseconds of startup is a trivial price against an escape that crosses tenants.

Choosing an ADK code executor

ADK treats execution as a pluggable strategy rather than something baked into the agent, which is exactly the seam you want: nothing about the agent changes when you move up the isolation ladder. Executors live under google.adk.code_executors and attach through the agent’s code_executor parameter.

from google.adk.agents import LlmAgent
from google.adk.code_executors import (
    BuiltInCodeExecutor,        # the model provider runs it, off your infra
    ContainerCodeExecutor,      # your Docker image, your limits
    VertexAiCodeExecutor,       # managed, sandboxed, stateful
    UnsafeLocalCodeExecutor,    # in-process. the name is the warning.
)

analyst = LlmAgent(
    name="analyst",
    model="gemini-2.0-flash",
    instruction="Analyze the attached dataset with Python.",
    code_executor=VertexAiCodeExecutor(),   # or ContainerCodeExecutor(image=...)
)

The trade-offs track the ladder. BuiltInCodeExecutor delegates to the model provider’s own execution tool: nothing runs on your infrastructure — the strongest possible statement about blast radius — but you inherit the provider’s library set and cannot reach your own data. ContainerCodeExecutor gives you the image, so you control dependencies, user, mounts, and network — and you own the hardening, because a default Docker configuration is not a sandbox. VertexAiCodeExecutor rents a managed sandbox with a persistent kernel. Whichever you pick, make the choice environment-dependent and fail startup loudly if the unsafe executor is selected outside development.

Resource limits: CPU, memory, wall clock, output size

Isolation stops the code reaching out; limits stop it consuming everything where it stands. Both get hit in practice — the first by attackers, the second, far more often, by an agent that wrote an accidental infinite loop.

LimitStarting pointWhat it prevents
CPU1–2 cores, hard-cappedMining, brute force, starving co-tenants
Memory512 MB–2 GB, OOM-kill on breachNode eviction from one allocation
Wall clock30–60 s per executionHangs, sleeps, slow-loris outbound connections
Processes (PIDs)64–256Fork bombs
Disk / tmpfs256 MB–1 GBFilling the node, hiding large staged payloads
Output size~64 KB, truncated with a markerContext blowout and token-cost attacks
Executions per session10–20Retry loops burning budget forever

Wall clock and output size are the two people forget. A timeout must kill the whole process group, or a spawned child outlives it. And output size is an attack surface, not hygiene: whatever the sandbox prints enters the model’s context and lands on your bill, so an injected print('A' * 10_000_000) is denial-of-wallet with no exploit required. Truncate at the boundary and append an [output truncated] marker so the model knows it sees a fragment.

Filesystem policy — read-only root, ephemeral scratch

The default posture is a read-only root filesystem with exactly one writable location: a small, size-capped tmpfs scratch directory that dies with the sandbox. Real work still fits — pandas needs somewhere to write a CSV, matplotlib somewhere to drop a PNG — while an attacker gets nowhere to install anything that survives.

Getting data in and out is where discipline breaks. The instinct is to bind-mount the directory the data already lives in, and that instinct is how a sandbox quietly becomes a shell on your host. Instead, copy the specific input files into scratch before execution and copy declared outputs back after, with an explicit allowlist of what may cross. In ADK the natural transport is the artifact service: the tool loads the artifact, writes its bytes into the sandbox, and saves any produced file back through ToolContext.

Three specifics worth hardcoding: run as a non-root user with a numeric UID that owns nothing on the host; drop all Linux capabilities and add none back, because analysis code needs zero of them; and set no-new-privileges, which makes setuid-binary escalation a non-event.

Scope the whole sandbox to one ADK session, destroyed at session end or after a short idle timeout, and never reuse one across users. Session scoping is what lets the iterative loop keep a loaded dataframe between turns without letting a poisoned sitecustomize.py written in one conversation lie in wait for the next user’s code — one conversation, one blast radius. If warm-start latency forces pooling, pool only empty sandboxes that have never executed anything.

Network egress is the exfiltration boundary

If the sandbox cannot open a connection, most attacks reduce to vandalism inside a container that is about to be deleted. That makes egress the highest-value control, and the default posture is simple: no network at all. Most analysis workloads genuinely do not need one, and a sandbox in an isolated network namespace with no route out is trivially auditable.

When code truly must fetch something, do not open the network — route it through an egress proxy that terminates the connection, checks the destination against an allowlist of exact hostnames, and logs every attempt. Allowlist by host and port, never by ‘anything but private ranges’, because attacker-controlled DNS collapses that distinction — and block DNS itself, since exfiltration through crafted subdomain lookups needs no HTTP at all.

The one destination to deny before all others is the cloud metadata endpoint at 169.254.169.254. A single unauthenticated HTTP GET there returns the node’s service-account token — the fastest path from ‘model wrote weird code’ to ‘attacker holds your cloud credentials’. Deny link-local, RFC1918, and loopback at the network layer, and page on any attempt — no legitimate analysis snippet has a reason to try.

Secrets the sandbox cannot reach

Every control above can fail. This one is the backstop, and it is the cheapest rule in the article: the sandbox has no credentials in it. No environment variables holding keys, no mounted service-account JSON, no cached token files, no ~/.config inherited from a base image, no ambient workload identity. If os.environ is worth nothing, the flagship injection payload — read the environment, POST it somewhere — returns an empty dict.

The pattern that makes this practical is brokering: privileged work happens in your trusted tool layer on the sandbox’s behalf, never inside it. Generated code calls a narrow, schema-validated function — ‘run this parameterized query against this dataset’ — and the tool, outside the boundary, resolves a short-lived scoped credential, performs the call, and hands back only the rows. The credential never enters the sandbox’s address space, never appears in a prompt, never lands in a transcript.

Scope those credentials as far down as the platform allows: one dataset, read-only, minutes of validity, bound to the session. Then verify empirically — run a canary session whose code dumps the environment and filesystem, and diff it against what you expect. It is a five-minute test that catches the base image someone helpfully added a key to.