ADK Java gives you six ways to put agents together: an LlmAgent that transfers to its sub-agents, an LlmAgent that calls another agent as a tool through AgentTool, the three workflow agents SequentialAgent, ParallelAgent and LoopAgent, and a custom agent that extends BaseAgent. All six compose, so almost any design can be built several ways. The expensive mistakes are not bugs in one pattern; they come from picking a pattern that puts the wrong party in charge of control flow.
Each pattern has its own deep dive on this site. This article is the layer above them: the questions that pick a pattern, what each choice costs in model calls, latency and testability, how state has to flow between agents for the choice to work, and one realistic assistant designed by walking the questions in order. Code uses the builder APIs in the google/adk-java repository and reads the model name from an environment variable, so substitute whatever model your project can call.
The one question that matters first
Every orchestration decision starts with: who decides what runs next, your code or the model? Workflow agents are code. A SequentialAgent runs its sub-agents in the order you listed them, a ParallelAgent runs them all at once, and a LoopAgent repeats them. None of the three makes a model call of its own; they are control structures written in Java. Transfer and AgentTool are the model: an LlmAgent reads the conversation and the descriptions of the agents it can reach, and chooses.
The difference is not stylistic. A code decision is deterministic, costs no tokens, and can be unit tested with an assertion. A model decision is probabilistic, costs at least one model call, and can only be evaluated statistically, with a labelled set of inputs and an accuracy number. So if you can write the decision as an if statement, do, and put the model in charge only where it needs language understanding.
| Building block | Who decides | Model calls it adds | What it gives you |
|---|---|---|---|
| SequentialAgent | code | none | fixed order, each stage sees earlier outputs through state |
| ParallelAgent | code | none | concurrent branches, events merged into one stream |
| LoopAgent | code, plus an escalate signal | none | repetition until escalate or maxIterations |
| Transfer (subAgents on an LlmAgent) | model | the routing call | hands the turn to a specialist |
| AgentTool | model | the call that picks the tool | a specialist answers one question, the caller continues |
| Custom BaseAgent | code you write | none of its own | any control flow: branch, halt, dynamic fan-out |
The decision tree
The diagram turns that first question and three follow-ups into a tree. Walk it for each part of your workflow: real assistants usually have a model-driven front door and deterministic pipelines behind it.
Code decides, steps independent: ParallelAgent. In the current source it subscribes each sub-agent's event stream on a scheduler and merges them, so branches really do overlap, and it gives each branch a name of the form parent.child so their conversation views stay separate. Fan-out pays when branches are slow and truly independent, such as several reviews of one document; see parallel fan-out in ADK Java.
Code decides, steps dependent, fixed count: SequentialAgent. Each stage writes its result to session state and the next stage reads it. It is the default for extraction and decision pipelines; see sequential agent chains.
Code decides, repeat until good enough: LoopAgent. It repeats its sub-agents until an event carries the escalate action or until maxIterations is reached. A draft-and-review pair is the classic use.
Model decides, specialist replies: transfer. List specialists as subAgents of an LlmAgent and the model can hand the current turn to one of them. Which agent gets the next message is a separate rule: in the current source the runner picks the agent that last replied only if it and every ancestor are LlmAgents that allow transfer to their parent; otherwise the next message starts at the root. So an LlmAgent specialist can keep a conversation, while a workflow target hands it back.
Model decides, caller keeps control: AgentTool. Wrap an agent with AgentTool.create(agent) and add it to the caller's tools. The caller invokes it like a function, gets back its answer and continues; the user never talks to the specialist directly. The trade-offs between the two delegation mechanisms are covered in hierarchical agents and supervisors.
What each choice costs
| Pattern | Model calls per turn (minimum) | Latency shape | Determinism | How you test it |
|---|---|---|---|---|
| Single LlmAgent with tools | 1 + one per tool round | sequential tool rounds | low | evals on final answers |
| SequentialAgent of N LlmAgents | N | sum of stages | order fixed, content not | unit test each stage, then the chain |
| ParallelAgent of N LlmAgents | N | slowest branch | order of events varies | per-branch tests plus a merge test |
| LoopAgent, K iterations of M agents | K x M | grows with iterations | iteration count varies | cap K; test exit and non-exit paths |
| Transfer | 1 routing call, then the specialist | adds one hop | low | routing accuracy on a labelled set |
| AgentTool | caller call + specialist calls + caller again | adds a round trip | low | tool-selection evals and specialist evals |
| Custom BaseAgent | whatever its children make | you control it | as high as your code | plain JUnit on the branching logic |
Two consequences are easy to miss. First, AgentTool is the most expensive way to delegate one question: the caller's model decides to call the tool, the specialist runs, and the caller's model runs again to use the result. Use it to combine specialists' answers, not to route. Second, a LoopAgent without maxIterations has no upper bound on cost: if the reviewer never escalates, it repeats until something else stops it. Always set the cap, and treat hitting it as an outcome to handle, not an error to ignore.
State is the contract between agents
Whatever pattern you pick, agents communicate through the session. The main mechanism is outputKey: when an LlmAgent finishes, ADK stores its final response text, excluding thoughts and events that only carry function calls, in session state under that key. Later instructions read it with a {key} placeholder. This is how a sequential chain passes results forward, and how the stages after a parallel fan-out see every branch's result; passing context between steps goes deeper.
- Parallel branches share one state. Branch names separate their conversation histories, not their state. Give each branch a distinct
outputKey; two branches writing the same key leave whichever finished last. - AgentTool returns a value, not a key. In the current source, the tool result is the specialist's last text wrapped as
{"result": ...}, or the structured output when the specialist has an output schema. State changes made inside the specialist are applied to the caller's state as well, so name keys as if they were global. - Transfer shares everything. The specialist sees the same session and history. That is convenient and also a leak: a specialist with fewer permissions reads everything the front door saw.
- Write down each key's owner. For every key, record which agent writes it, which agents read it, and its format. Most multi-agent bugs are two agents disagreeing about one of those three.
When none of the shapes fits: a custom agent
The workflow agents cannot branch on a value. "If validation failed, stop; if the amount is over the limit, go to the approval path; otherwise continue" is not a sequence, a fan-out or a loop, and handing that decision to a model would make a deterministic rule probabilistic. That is what extending BaseAgent is for: you implement runAsyncImpl, read state from the invocation context, and choose which child to run.
// imports: com.google.adk.agents.{BaseAgent, InvocationContext}, com.google.adk.events.Event,
// io.reactivex.rxjava3.core.Flowable
public final class AmountRouter extends BaseAgent {
private final BaseAgent autoApprove;
private final BaseAgent needsApproval;
public AmountRouter(BaseAgent autoApprove, BaseAgent needsApproval) {
super("amount_router", "Routes an expense by amount; no model call.",
List.of(autoApprove, needsApproval), List.of(), List.of());
this.autoApprove = autoApprove;
this.needsApproval = needsApproval;
}
static boolean overLimit(Object receiptJson) { // pure: unit test it directly
return new BigDecimal(Receipts.amount(receiptJson)).compareTo(new BigDecimal("500")) > 0;
}
@Override
protected Flowable<Event> runAsyncImpl(InvocationContext ctx) {
return Flowable.defer(() ->
(overLimit(ctx.session().state().get("receipt")) ? needsApproval : autoApprove)
.runAsync(ctx));
}
@Override
protected Flowable<Event> runLiveImpl(InvocationContext ctx) {
return Flowable.error(new UnsupportedOperationException("live mode not supported"));
}
}Receipts.amount stands for your own JSON parsing. The decision lives in a static method that plain JUnit can test, and the agent itself is three lines of plumbing. Custom agents are also how you halt a chain early and fan out over a list whose length is only known at run time; extending the ADK Java runtime covers the contract in detail.
Worked example: an expense assistant
Requirements: employees chat with an assistant. They either ask policy questions or submit a receipt. A submitted receipt is extracted, checked against the spending limit and for duplicates, and answered with a clear, policy-compliant message. Walk the tree for each part.
- Front door. Whether a message is a question or a submission needs language understanding: model decides. A submission is handed to the expense flow by transfer; because that flow is a workflow agent, the next message returns to the front door, which is what we want. Policy questions are answered by the front door itself using a policy specialist behind
AgentTool, because the front door keeps the conversation. - Submission flow. Extract, then check, then respond, always in that order:
SequentialAgent. - Checks. The limit check and the duplicate check do not depend on each other:
ParallelAgentwith distinct output keys. - Response. Draft, review, redraft until compliant, at most three times:
LoopAgentwithExitLoopTool.INSTANCEon the reviewer.
// imports: com.google.adk.agents.{BaseAgent, LlmAgent, LoopAgent, ParallelAgent, SequentialAgent},
// com.google.adk.tools.{AgentTool, ExitLoopTool}
static final String MODEL = System.getenv("ADK_MODEL");
static LlmAgent agent(String name, String instruction, String outputKey) {
return LlmAgent.builder().name(name).model(MODEL)
.instruction(instruction).outputKey(outputKey).build();
}
static BaseAgent build() {
LlmAgent extract = agent("extract",
"Extract merchant, date, amount and currency from the receipt. Reply with JSON only.", "receipt");
ParallelAgent checks = ParallelAgent.builder().name("checks").subAgents(
agent("limit_check", "Receipt: {receipt}. Is it within the travel limit? Reply OK or OVER with a reason.", "limit_check"),
agent("dupe_check", "Receipt: {receipt}. Use the history tool; reply UNIQUE or DUPLICATE.", "dupe_check"))
.build(); // history tool omitted for brevity
LlmAgent writer = agent("writer",
"Write the reply to the employee. Receipt: {receipt}. Limit: {limit_check}. Duplicates: {dupe_check}. "
+ "If a reviewer asked for changes earlier in the conversation, apply them.", "draft");
LlmAgent reviewer = LlmAgent.builder().name("reviewer").model(MODEL)
.instruction("Draft: {draft}. If it states the outcome, the reason and the next step, call exitLoop. "
+ "Otherwise list the exact changes needed.")
.tools(ExitLoopTool.INSTANCE).build();
LoopAgent respond = LoopAgent.builder().name("respond")
.subAgents(writer, reviewer).maxIterations(3).build();
SequentialAgent expenseFlow = SequentialAgent.builder().name("expense_flow")
.description("Processes a submitted receipt: extraction, checks and a reply.")
.subAgents(extract, checks, respond).build();
LlmAgent policy = agent("policy_qa", "Answer questions about the expense policy, quoting the rule.", "policy_answer");
return LlmAgent.builder().name("front_desk").model(MODEL)
.instruction("If the user submits a receipt, transfer to expense_flow. "
+ "For policy questions, call policy_qa and answer in your own words.")
.subAgents(expenseFlow)
.tools(AgentTool.create(policy))
.build();
}Count the cost of a submission: one front-door routing call, one extraction, two parallel checks, then two to six calls in the loop. That is six to ten model calls, with the checks costing one call's latency rather than two. A policy question costs three: the front door, the specialist, and the front door again. If the AmountRouter above were added after the checks, it would add no model calls at all, which is exactly why that decision belongs in code.
Failure modes
- A model doing a workflow's job. A front-door agent told to "first extract, then check, then reply" will sometimes skip or reorder steps. If the order is fixed, encode it in a
SequentialAgent. - Transfer ping-pong. Two agents that can transfer to each other bounce a message back and forth. Use
disallowTransferToPeers(true)anddisallowTransferToParent(true)where a specialist should finish its job, and write descriptions that do not overlap; see agent routing for how descriptions drive routing. - Unbounded loops. A reviewer prompt that never calls
exitLoopon borderline drafts turns into maximum cost on every request. SetmaxIterations, track how often it is reached, and treat a high rate as a prompt bug. - Clobbered state. Parallel branches or nested agents writing the same key. Keep a key registry and assert distinct keys in a unit test that walks the agent tree.
- Missing placeholders. An instruction that reads
{limit_check}before any agent has written it can fail at run time; check how your version treats a missing key. Order stages so every key is written before it is read, or seed it when the session is created.
What to do next
- List each step in your workflow and mark who should decide what runs after it: code or model.
- Walk the decision tree for each part and sketch the agent tree before writing any prompts.
- Write a state-key table: key, writer, readers, format. Check that parallel branches use distinct keys.
- Count the minimum model calls per request type from the cost table and compare with your latency and cost budget.
- Move every rule you can express as an
ifinto a workflow agent or a customBaseAgent, with a JUnit test for the rule. - Build a labelled set of user messages for each model-driven decision (transfer or tool choice) and measure routing accuracy before launch.
- Set
maxIterationson everyLoopAgentand alert when runs hit it.