A common description of the ADK Java runtime says that every session gets its own virtual thread, so blocking code is always free. That is not how the library works. ADK Java targets Java 17, which has no virtual threads as a standard feature, and the runtime does not start a thread per session at all. It hands you a lazy reactive stream, and the thread that subscribes to that stream does most of the work, including waiting for the model.
That design is flexible, and it puts the decision with you. This article traces which thread runs each part of an agent turn, explains the knobs that move work elsewhere, and shows how to host the runtime on a servlet container, an event-loop server and a command-line job without stalling anything. The source references are to google/adk-java at release v1.10.1, from September 2026; check them against the version you run, because this part of the code has changed between releases.
The runtime owns no threads
Runner.runAsync(...) returns a Flowable<Event> from RxJava 3 built with Flowable.defer. It is cold: calling runAsync does nothing but assemble a pipeline. Loading the session, running the agent's request processors, calling the model, dispatching tools and appending events all happen when something subscribes, and, unless an operator moves the work elsewhere, they happen on the subscribing thread, one step after another.
So the right question is never how many threads ADK uses. It is which thread subscribes, and which operators in the pipeline move work off it. The per-stage breakdown of one loop iteration is covered in the execution loop anatomy; this article covers the switches that change the answer and how to pick them for your host.
The model call blocks the subscriber
The Gemini model class calls the asynchronous API of the Google Gen AI Java client, which returns a CompletableFuture, and wraps it with Flowable.fromFuture. In RxJava 3 that operator waits for the future with a blocking get() on the subscribing thread. The network I/O happens on the HTTP client's dispatcher threads, but your thread is parked for the whole round trip. For streaming responses the stream is iterated on the same subscribing thread, so it is parked between chunks as well. For non-streaming calls, the conversion to the ADK response type runs through thenApplyAsync, which uses the common fork-join pool unless an executor is given.
Two consequences follow. First, a model call of several seconds holds whichever thread subscribed for those seconds, once per loop iteration. On a virtual thread that costs almost nothing; on a Netty event-loop thread it stalls every connection that loop serves. Second, the HTTP client that ADK builds for Gemini sets connect, read and write timeouts to zero, which in OkHttp means no timeout. A stuck connection therefore holds the subscriber indefinitely. Two more facts from RxJava 3's own documentation make this sharper: because the wait happens inside subscribe(), a downstream timeout can signal an error but cannot unpark the thread that is waiting, and cancelling a fromFuture stream does not cancel the future. The hosting examples below deal with both.
Gemini's builder accepts httpExecutorService(...), which sets the executor for the HTTP dispatcher. Without it the model uses a shared client. The builder's documentation suggests a daemon executor for standalone and command-line programs so that the JVM can exit when work is done, and a container-managed executor in managed environments. It is ignored when you pass your own apiClient.
Tool execution modes
When one model response asks for several function calls, RunConfig.toolExecutionMode decides how they run. There are four values, and the difference between the two parallel ones is the most important thread-model fact in the library.
| Mode | Operator | Where tool bodies run | Blocking tools |
|---|---|---|---|
SEQUENTIAL, or any single call | concatMapMaybe | Subscriber thread, one at a time | Sequential |
PARALLEL | concatMapEager | Subscriber thread; all subscribed eagerly | Still sequential; only async tools overlap |
NONE (the default) | Treated as PARALLEL | As PARALLEL | As PARALLEL |
PARALLEL_SUBSCRIBE | concatMapEager with subscribeOn | Agent executor, else Schedulers.io() | Concurrent |
The default is parallel in name only for blocking tools. A plain Java method registered with FunctionTool is invoked reflectively inside Maybe.defer, which runs on whichever thread subscribes to it, so under PARALLEL three blocking HTTP tools run one after the other on the caller's thread. Tools that return a Single or Maybe backed by real asynchronous I/O do overlap, because subscribing to them returns immediately. PARALLEL_SUBSCRIBE adds subscribeOn to each call, so blocking bodies run on separate workers. In every mode responses stay in call order, which the runtime needs when it merges them into one event.
public class OrderTools {
// Blocking tool: fine on a virtual thread or a worker, harmful on an event loop.
public static Map<String, Object> getOrder(@Schema(name = "orderId") String orderId) {
return orderClient.fetch(orderId); // plain blocking HTTP call
}
// Non-blocking tool: returns immediately, so PARALLEL already overlaps it with siblings.
public static Single<Map<String, Object>> getShipment(@Schema(name = "orderId") String orderId) {
return Single.fromCompletionStage(shippingClient.fetchAsync(orderId))
.timeout(5, TimeUnit.SECONDS);
}
}PARALLEL_SUBSCRIBE has a cost: tools now run concurrently with each other, so any shared state they touch must be thread-safe, and a tool that assumed it ran on the request thread loses whatever thread-local context that thread carried. Per-tool deadlines belong in the tool; tool timeout handling covers the patterns.
The agent executor and the other knobs
LlmAgent.builder().executor(Executor) sets the worker for that agent. In the code checked, it is used in two places: as the scheduler for PARALLEL_SUBSCRIBE tool calls, and as the scheduler the live, bidirectional-streaming flow observes on when it sends queued requests. When it is not set, both fall back to Schedulers.io(), RxJava's cached pool, which grows without an upper bound. Under load, blocking tools on that scheduler can create hundreds of platform threads.
import com.google.adk.agents.LlmAgent;
import com.google.adk.agents.RunConfig;
import com.google.adk.models.Gemini;
import com.google.adk.runner.InMemoryRunner;
import java.util.concurrent.Executors;
// JDK 21+: one virtual thread per task. On JDK 17 use a bounded platform pool instead.
var toolWorkers = Executors.newVirtualThreadPerTaskExecutor();
Gemini model = Gemini.builder()
.modelName("gemini-2.5-flash")
.apiKey(System.getenv("GOOGLE_API_KEY"))
.httpExecutorService(Executors.newFixedThreadPool(16)) // OkHttp dispatcher threads
.build();
LlmAgent agent = LlmAgent.builder()
.name("orders")
.model(model)
.instruction("Answer order questions using the tools.")
.tools(/* FunctionTool.create(...) */)
.executor(toolWorkers) // used by PARALLEL_SUBSCRIBE and the live send loop
.build();
RunConfig runConfig = RunConfig.builder()
.toolExecutionMode(RunConfig.ToolExecutionMode.PARALLEL_SUBSCRIBE)
.maxLlmCalls(20) // default is 500
.build();
var runner = new InMemoryRunner(agent);The model name above is only an example; use whichever model your project is configured for. maxLlmCalls is not a thread setting, but it bounds how many times one invocation can call the model, and therefore how long one subscriber can be held; the default of 500 is far above what a well-behaved turn needs.
Virtual threads and pinning, by JDK version
Virtual threads, final in JDK 21, are the natural subscriber for this runtime: parking a virtual thread on a future releases its carrier thread, so thousands of turns can wait on the model at once. The general case for them in agent code is made in virtual threads for Java agents. Because ADK targets Java 17, it cannot create them for you; you supply them, either as the thread that subscribes or through executor(...).
Pinning is where the JDK version matters. On JDK 21 to 23, a virtual thread that blocks while inside a synchronized block or method stays mounted on its carrier, so a few such waits can occupy every carrier and stall unrelated work. JEP 491, delivered in JDK 24, removed that limitation for synchronized. Pinning still happens when a virtual thread blocks while a native frame is on its stack, such as a JNI call. Libraries you call from tools, including JDBC drivers and older HTTP clients, may still use synchronized internally, so on JDK 21 the practical rule is to measure rather than assume: JDK Flight Recorder's jdk.VirtualThreadPinned event reports pinned waits with stack traces.
Hosting: pick the thread that pays
The same agent needs a different subscription strategy depending on the server around it.
Scheduler agentScheduler = Schedulers.from(Executors.newVirtualThreadPerTaskExecutor());
// 1. Blocking servlet (Spring MVC). fromFuture blocks inside subscribe(), so without
// subscribeOn the timeout would fire while this thread stays stuck in Future.get().
List<Event> events = runner.runAsync(userId, sessionId, message, runConfig)
.subscribeOn(agentScheduler)
.timeout(90, TimeUnit.SECONDS) // frees the caller; does not stop the request
.toList()
.blockingGet();
// 2. Event-loop server (WebFlux / Netty / Vert.x): never subscribe on the I/O thread.
Flowable<Event> stream = runner.runAsync(userId, sessionId, message, runConfig)
.subscribeOn(agentScheduler) // blocking waits land here, not on the event loop
.timeout(90, TimeUnit.SECONDS);
// adapt to the framework, e.g. Flux.from(stream) for Reactor-based servers
// 3. CLI or batch job: the main thread may block; let the JVM exit afterwards.
runner.runAsync(userId, sessionId, message, runConfig)
.blockingForEach(e -> System.out.println(e.stringifyContent()));
// 4. A real bound on the model call: pass your own Gen AI client with an HTTP timeout
// (HttpOptions.timeout is in milliseconds). ADK's zero-timeout client is not used then.
Client client = Client.builder()
.apiKey(System.getenv("GOOGLE_API_KEY"))
.httpOptions(HttpOptions.builder().timeout(60_000).build())
.build();
Gemini boundedModel = new Gemini("gemini-2.5-flash", client);On a blocking servlet stack, the request thread may simply wait. With virtual threads enabled in the container that is the simplest and cheapest arrangement; with a fixed platform pool of, say, 200 threads, it caps concurrency at 200 turns, and the pool's threads spend nearly all their time parked on the model. On an event-loop server the rule is absolute: subscribe on a dedicated scheduler, never on the I/O thread, because every blocking wait described above would stall the loop. In a command-line job, blocking the main thread is fine; set a daemon HTTP executor so the process exits when it finishes. Whatever the host, bound the turn twice: a stream-level timeout after subscribeOn, which frees the caller but leaves the in-flight request running on its worker, and an HTTP timeout on a client you supply to Gemini, which actually ends a stuck model call.
Context propagation and cancellation
Moving work between threads breaks anything stored in thread-locals. ADK handles its own tracing: Runner, the LLM flow and the function dispatcher capture the OpenTelemetry Context.current() when the pipeline is assembled and reapply it inside the deferred work, so spans nest correctly even when tools run on IO workers. Your own thread-locals, such as SLF4J MDC values, security contexts and request-scoped beans, are not propagated. Pass what tools need through the tool context or session state, or wrap the executor you give ADK so that it copies the values you need onto each task.
Cancellation is disposal of the subscription. Disposing stops events from reaching your subscriber, but it does not cancel the model future, and a blocking tool body stops only if it responds to interruption and the scheduler interrupts on disposal. Test it in your version by disposing mid-turn and watching the HTTP client and tool logs. For graceful draining of in-flight turns at shutdown, see graceful shutdown.
Worked example: three tools, three modes
Suppose the model asks for three tools in one response, each a blocking HTTP call of 400 ms, and the subscriber is a virtual thread. Under the default mode, the three bodies run one after another on that virtual thread: 1.2 seconds of tool time, with no platform thread held. Rewrite the three tools to return Single from an asynchronous client and the default mode overlaps them, about 400 ms. Keep them blocking but switch to PARALLEL_SUBSCRIBE with a virtual-thread executor, and they also overlap, about 400 ms, at the cost of thread-safety discipline. Put the same blocking tools on a Netty I/O thread in the default mode and that loop is frozen for 1.2 seconds per turn, and every other connection on it waits too.
| Failure mode | Symptom | Fix |
|---|---|---|
| Subscribing on an event-loop thread | Unrelated requests time out during agent turns | subscribeOn a dedicated scheduler |
Assuming PARALLEL overlaps blocking tools | Turn latency equals the sum of tool latencies | PARALLEL_SUBSCRIBE or async tools |
No timeouts, or a stream timeout without subscribeOn | Turns that never finish; threads held forever | subscribeOn then .timeout(...), plus an HTTP timeout on your own client |
Unbounded Schedulers.io() | Thread count climbs under load | Set executor(...) to a bounded or virtual-thread executor |
| Pinning on JDK 21 to 23 | Carriers exhausted, throughput collapses | Upgrade to 24+, or find pins with JFR |
| Lost MDC or security context in tools | Logs without request IDs, authorization failures | Pass data explicitly or wrap the executor |
What to do next
- Find every place your code subscribes to
runAsyncand write down which thread that is. - If any of them is an event-loop or small shared pool, add
subscribeOnwith a dedicated scheduler. - Add
subscribeOnplus a stream timeout to every invocation, and supply a Gen AI client with an HTTP timeout. - Classify each tool as blocking or asynchronous; choose
PARALLEL_SUBSCRIBEor rewrite hot tools to returnSingle. - Give each agent an explicit executor so that nothing depends on the unbounded IO scheduler.
- On JDK 21 to 23, record a JFR session under load and check for
jdk.VirtualThreadPinnedevents. - Dispose a subscription mid-turn in a test and confirm what actually stops.