A common description of the ADK Java runtime says that every session gets its own virtual thread, so blocking code is always free. That is not how the library works. ADK Java targets Java 17, which has no virtual threads as a standard feature, and the runtime does not start a thread per session at all. It hands you a lazy reactive stream, and the thread that subscribes to that stream does most of the work, including waiting for the model.

That design is flexible, and it puts the decision with you. This article traces which thread runs each part of an agent turn, explains the knobs that move work elsewhere, and shows how to host the runtime on a servlet container, an event-loop server and a command-line job without stalling anything. The source references are to google/adk-java at release v1.10.1, from September 2026; check them against the version you run, because this part of the code has changed between releases.

Advertisement

The runtime owns no threads

Runner.runAsync(...) returns a Flowable<Event> from RxJava 3 built with Flowable.defer. It is cold: calling runAsync does nothing but assemble a pipeline. Loading the session, running the agent's request processors, calling the model, dispatching tools and appending events all happen when something subscribes, and, unless an operator moves the work elsewhere, they happen on the subscribing thread, one step after another.

So the right question is never how many threads ADK uses. It is which thread subscribes, and which operators in the pipeline move work off it. The per-stage breakdown of one loop iteration is covered in the execution loop anatomy; this article covers the switches that change the answer and how to pick them for your host.

Caller threadsubscribes to runAsync(...)Runner / LlmAgent flowcold Flowable, runs on subscribersubscribeOTel Contextcaptured at assemblyModel callFlowable.fromFuture: get() blocksOkHttp dispatchershared or httpExecutorServicesocket I/OSEQUENTIAL / 1 callconcatMapMaybe, caller threadPARALLEL (default)concatMapEager, caller threadPARALLEL_SUBSCRIBEsubscribeOn(executor or io())Blocking tool bodies run one after another unless PARALLEL_SUBSCRIBE moves them to workersResult: the thread that subscribes pays for every blocking wait in the turnchoose it deliberately: virtual thread, bounded worker pool, or a request thread you can afford to park
Where time is spent in one turn. The subscriber drives the flow; the model call parks it on a future; tools run on the caller thread unless PARALLEL_SUBSCRIBE moves them to the agent executor or RxJava's IO scheduler.

The model call blocks the subscriber

The Gemini model class calls the asynchronous API of the Google Gen AI Java client, which returns a CompletableFuture, and wraps it with Flowable.fromFuture. In RxJava 3 that operator waits for the future with a blocking get() on the subscribing thread. The network I/O happens on the HTTP client's dispatcher threads, but your thread is parked for the whole round trip. For streaming responses the stream is iterated on the same subscribing thread, so it is parked between chunks as well. For non-streaming calls, the conversion to the ADK response type runs through thenApplyAsync, which uses the common fork-join pool unless an executor is given.

Two consequences follow. First, a model call of several seconds holds whichever thread subscribed for those seconds, once per loop iteration. On a virtual thread that costs almost nothing; on a Netty event-loop thread it stalls every connection that loop serves. Second, the HTTP client that ADK builds for Gemini sets connect, read and write timeouts to zero, which in OkHttp means no timeout. A stuck connection therefore holds the subscriber indefinitely. Two more facts from RxJava 3's own documentation make this sharper: because the wait happens inside subscribe(), a downstream timeout can signal an error but cannot unpark the thread that is waiting, and cancelling a fromFuture stream does not cancel the future. The hosting examples below deal with both.

Gemini's builder accepts httpExecutorService(...), which sets the executor for the HTTP dispatcher. Without it the model uses a shared client. The builder's documentation suggests a daemon executor for standalone and command-line programs so that the JVM can exit when work is done, and a container-managed executor in managed environments. It is ignored when you pass your own apiClient.

Advertisement

Tool execution modes

When one model response asks for several function calls, RunConfig.toolExecutionMode decides how they run. There are four values, and the difference between the two parallel ones is the most important thread-model fact in the library.

ModeOperatorWhere tool bodies runBlocking tools
SEQUENTIAL, or any single callconcatMapMaybeSubscriber thread, one at a timeSequential
PARALLELconcatMapEagerSubscriber thread; all subscribed eagerlyStill sequential; only async tools overlap
NONE (the default)Treated as PARALLELAs PARALLELAs PARALLEL
PARALLEL_SUBSCRIBEconcatMapEager with subscribeOnAgent executor, else Schedulers.io()Concurrent

The default is parallel in name only for blocking tools. A plain Java method registered with FunctionTool is invoked reflectively inside Maybe.defer, which runs on whichever thread subscribes to it, so under PARALLEL three blocking HTTP tools run one after the other on the caller's thread. Tools that return a Single or Maybe backed by real asynchronous I/O do overlap, because subscribing to them returns immediately. PARALLEL_SUBSCRIBE adds subscribeOn to each call, so blocking bodies run on separate workers. In every mode responses stay in call order, which the runtime needs when it merges them into one event.

public class OrderTools {
  // Blocking tool: fine on a virtual thread or a worker, harmful on an event loop.
  public static Map<String, Object> getOrder(@Schema(name = "orderId") String orderId) {
    return orderClient.fetch(orderId);             // plain blocking HTTP call
  }

  // Non-blocking tool: returns immediately, so PARALLEL already overlaps it with siblings.
  public static Single<Map<String, Object>> getShipment(@Schema(name = "orderId") String orderId) {
    return Single.fromCompletionStage(shippingClient.fetchAsync(orderId))
        .timeout(5, TimeUnit.SECONDS);
  }
}

PARALLEL_SUBSCRIBE has a cost: tools now run concurrently with each other, so any shared state they touch must be thread-safe, and a tool that assumed it ran on the request thread loses whatever thread-local context that thread carried. Per-tool deadlines belong in the tool; tool timeout handling covers the patterns.

The agent executor and the other knobs

LlmAgent.builder().executor(Executor) sets the worker for that agent. In the code checked, it is used in two places: as the scheduler for PARALLEL_SUBSCRIBE tool calls, and as the scheduler the live, bidirectional-streaming flow observes on when it sends queued requests. When it is not set, both fall back to Schedulers.io(), RxJava's cached pool, which grows without an upper bound. Under load, blocking tools on that scheduler can create hundreds of platform threads.

import com.google.adk.agents.LlmAgent;
import com.google.adk.agents.RunConfig;
import com.google.adk.models.Gemini;
import com.google.adk.runner.InMemoryRunner;
import java.util.concurrent.Executors;

// JDK 21+: one virtual thread per task. On JDK 17 use a bounded platform pool instead.
var toolWorkers = Executors.newVirtualThreadPerTaskExecutor();

Gemini model = Gemini.builder()
    .modelName("gemini-2.5-flash")
    .apiKey(System.getenv("GOOGLE_API_KEY"))
    .httpExecutorService(Executors.newFixedThreadPool(16))   // OkHttp dispatcher threads
    .build();

LlmAgent agent = LlmAgent.builder()
    .name("orders")
    .model(model)
    .instruction("Answer order questions using the tools.")
    .tools(/* FunctionTool.create(...) */)
    .executor(toolWorkers)            // used by PARALLEL_SUBSCRIBE and the live send loop
    .build();

RunConfig runConfig = RunConfig.builder()
    .toolExecutionMode(RunConfig.ToolExecutionMode.PARALLEL_SUBSCRIBE)
    .maxLlmCalls(20)                  // default is 500
    .build();

var runner = new InMemoryRunner(agent);

The model name above is only an example; use whichever model your project is configured for. maxLlmCalls is not a thread setting, but it bounds how many times one invocation can call the model, and therefore how long one subscriber can be held; the default of 500 is far above what a well-behaved turn needs.

Virtual threads and pinning, by JDK version

Virtual threads, final in JDK 21, are the natural subscriber for this runtime: parking a virtual thread on a future releases its carrier thread, so thousands of turns can wait on the model at once. The general case for them in agent code is made in virtual threads for Java agents. Because ADK targets Java 17, it cannot create them for you; you supply them, either as the thread that subscribes or through executor(...).

Pinning is where the JDK version matters. On JDK 21 to 23, a virtual thread that blocks while inside a synchronized block or method stays mounted on its carrier, so a few such waits can occupy every carrier and stall unrelated work. JEP 491, delivered in JDK 24, removed that limitation for synchronized. Pinning still happens when a virtual thread blocks while a native frame is on its stack, such as a JNI call. Libraries you call from tools, including JDBC drivers and older HTTP clients, may still use synchronized internally, so on JDK 21 the practical rule is to measure rather than assume: JDK Flight Recorder's jdk.VirtualThreadPinned event reports pinned waits with stack traces.

Hosting: pick the thread that pays

The same agent needs a different subscription strategy depending on the server around it.

Scheduler agentScheduler = Schedulers.from(Executors.newVirtualThreadPerTaskExecutor());

// 1. Blocking servlet (Spring MVC). fromFuture blocks inside subscribe(), so without
// subscribeOn the timeout would fire while this thread stays stuck in Future.get().
List<Event> events = runner.runAsync(userId, sessionId, message, runConfig)
    .subscribeOn(agentScheduler)
    .timeout(90, TimeUnit.SECONDS)            // frees the caller; does not stop the request
    .toList()
    .blockingGet();

// 2. Event-loop server (WebFlux / Netty / Vert.x): never subscribe on the I/O thread.
Flowable<Event> stream = runner.runAsync(userId, sessionId, message, runConfig)
    .subscribeOn(agentScheduler)              // blocking waits land here, not on the event loop
    .timeout(90, TimeUnit.SECONDS);
// adapt to the framework, e.g. Flux.from(stream) for Reactor-based servers

// 3. CLI or batch job: the main thread may block; let the JVM exit afterwards.
runner.runAsync(userId, sessionId, message, runConfig)
    .blockingForEach(e -> System.out.println(e.stringifyContent()));

// 4. A real bound on the model call: pass your own Gen AI client with an HTTP timeout
// (HttpOptions.timeout is in milliseconds). ADK's zero-timeout client is not used then.
Client client = Client.builder()
    .apiKey(System.getenv("GOOGLE_API_KEY"))
    .httpOptions(HttpOptions.builder().timeout(60_000).build())
    .build();
Gemini boundedModel = new Gemini("gemini-2.5-flash", client);

On a blocking servlet stack, the request thread may simply wait. With virtual threads enabled in the container that is the simplest and cheapest arrangement; with a fixed platform pool of, say, 200 threads, it caps concurrency at 200 turns, and the pool's threads spend nearly all their time parked on the model. On an event-loop server the rule is absolute: subscribe on a dedicated scheduler, never on the I/O thread, because every blocking wait described above would stall the loop. In a command-line job, blocking the main thread is fine; set a daemon HTTP executor so the process exits when it finishes. Whatever the host, bound the turn twice: a stream-level timeout after subscribeOn, which frees the caller but leaves the in-flight request running on its worker, and an HTTP timeout on a client you supply to Gemini, which actually ends a stuck model call.

Context propagation and cancellation

Moving work between threads breaks anything stored in thread-locals. ADK handles its own tracing: Runner, the LLM flow and the function dispatcher capture the OpenTelemetry Context.current() when the pipeline is assembled and reapply it inside the deferred work, so spans nest correctly even when tools run on IO workers. Your own thread-locals, such as SLF4J MDC values, security contexts and request-scoped beans, are not propagated. Pass what tools need through the tool context or session state, or wrap the executor you give ADK so that it copies the values you need onto each task.

Cancellation is disposal of the subscription. Disposing stops events from reaching your subscriber, but it does not cancel the model future, and a blocking tool body stops only if it responds to interruption and the scheduler interrupts on disposal. Test it in your version by disposing mid-turn and watching the HTTP client and tool logs. For graceful draining of in-flight turns at shutdown, see graceful shutdown.

Worked example: three tools, three modes

Suppose the model asks for three tools in one response, each a blocking HTTP call of 400 ms, and the subscriber is a virtual thread. Under the default mode, the three bodies run one after another on that virtual thread: 1.2 seconds of tool time, with no platform thread held. Rewrite the three tools to return Single from an asynchronous client and the default mode overlaps them, about 400 ms. Keep them blocking but switch to PARALLEL_SUBSCRIBE with a virtual-thread executor, and they also overlap, about 400 ms, at the cost of thread-safety discipline. Put the same blocking tools on a Netty I/O thread in the default mode and that loop is frozen for 1.2 seconds per turn, and every other connection on it waits too.

Failure modeSymptomFix
Subscribing on an event-loop threadUnrelated requests time out during agent turnssubscribeOn a dedicated scheduler
Assuming PARALLEL overlaps blocking toolsTurn latency equals the sum of tool latenciesPARALLEL_SUBSCRIBE or async tools
No timeouts, or a stream timeout without subscribeOnTurns that never finish; threads held foreversubscribeOn then .timeout(...), plus an HTTP timeout on your own client
Unbounded Schedulers.io()Thread count climbs under loadSet executor(...) to a bounded or virtual-thread executor
Pinning on JDK 21 to 23Carriers exhausted, throughput collapsesUpgrade to 24+, or find pins with JFR
Lost MDC or security context in toolsLogs without request IDs, authorization failuresPass data explicitly or wrap the executor

What to do next

  1. Find every place your code subscribes to runAsync and write down which thread that is.
  2. If any of them is an event-loop or small shared pool, add subscribeOn with a dedicated scheduler.
  3. Add subscribeOn plus a stream timeout to every invocation, and supply a Gen AI client with an HTTP timeout.
  4. Classify each tool as blocking or asynchronous; choose PARALLEL_SUBSCRIBE or rewrite hot tools to return Single.
  5. Give each agent an explicit executor so that nothing depends on the unbounded IO scheduler.
  6. On JDK 21 to 23, record a JFR session under load and check for jdk.VirtualThreadPinned events.
  7. Dispose a subscription mid-turn in a test and confirm what actually stops.
Key takeaway: ADK Java does not give each session a thread. runAsync returns a cold RxJava stream, and the thread that subscribes runs the turn and blocks on every model call. Tools run on that same thread unless you choose PARALLEL_SUBSCRIBE, which moves them to the agent's executor or the unbounded IO scheduler. Subscribe from a virtual thread or a dedicated scheduler, never from an event loop, bound turns with a stream timeout after subscribeOn and an HTTP timeout on your own model client, because ADK's default client has none, give agents explicit executors, and check pinning if you are on JDK 21 to 23.