An agent workflow that waits for a human, a batch job or a payment provider cannot hold a thread open for two days, and it cannot assume the process that started the work is the one that finishes it. Resumability is ADK's answer: the workflow pauses at a well-defined point, the request ends, and a later request picks up from where it stopped instead of re-running everything from the first agent.

ADK Java's support is real but moving quickly, and the released version and the main branch differ in important ways. This article explains the model from first principles, shows working code for the released behaviour, describes what the main branch adds, and covers the operational rules that apply to every version: persistent sessions, idempotent tools and not changing the workflow while it is paused. Every class and method named here was checked against the google/adk-java repository on 2026-10-02 (release v1.10.1 of 2026-09-18 and the main branch). Treat anything marked main-only as subject to change.

Advertisement

The model: invocations, events and a pause point

In ADK, one call to the runner is an invocation. The runner looks up the session, appends your message as an event, runs the root agent, and streams back every event the agents produce: model text, function calls, function responses, state deltas. The session service persists those events. The event log is the durable record of what happened, and it is the raw material for resuming (the runtime execution loop walks through one invocation in detail).

A long-running tool is a function tool whose result is not available when the function returns. In Java it is a LongRunningFunctionTool: the method starts the work (opens an approval ticket, submits a job) and returns something like a pending status. The event carrying that function call lists the call's id in longRunningToolIds. With resumability enabled, that marker is the pause point: the workflow stops instead of continuing to the next agent, and the invocation ends.

Resuming means sending, in a new request, the function response that answers the paused call. The runner must then work out which agent should continue and which agents already finished, so it does not re-run them. How it works that out is exactly where the released version and main differ.

Pause and resume across two requests (SequentialAgent, resumable app)Request 1runAsync(user msg)draft_agentruns, emits textapproval_agentcalls long-running toolPausepending call idSession event log (persistent session service)user msg, draft text, function call with longRunningToolIds... minutes or days pass; the process may restart ...Request 2runAsync(function response)Runner routesto top SequentialAgentapproval_agentresumes, reads responsepublish_agentruns nexthistory readdraft_agent is not re-run: the resume point is found from the log (v1.10.1) or from agentState checkpoints (main).Any tool that was mid-flight when the process died may run again: resumption is at-least-once.
Two requests, one workflow. The first ends at the long-running call; the second delivers the answer, and the runner re-enters the sequence at the agent that made the call.

What is released and what is on main

In v1.10.1, ResumabilityConfig exists with a resumable flag, and both the class and App.Builder.resumabilityConfig are annotated @Deprecated with a note that says it is a partial feature: only event-reconstruction pause and resume for SequentialAgent is implemented, and full session resumability (persisted agent state, durable resume, other workflow agents) is not yet available. The deprecation is a warning about completeness, not a removal notice; the note says the same config will drive full resumability.

On main, the annotation is @Experimental instead, and the javadoc describes checkpointing agent state as the workflow runs, with resume being best-effort and at-least-once. The v1.10.0 changelog already lists EventActions.agentState for session-resumability checkpoints, so the field ships in the release even though the released SequentialAgent does not write it.

Capabilityv1.10.1 (released)main (unreleased, 2026-10-02)
ResumabilityConfig.resumableyes, @Deprecated (partial)yes, @Experimental
SequentialAgent pause on long-running callyes, resume point rebuilt from eventsyes, checkpoint before each sub-agent
LoopAgent, ParallelAgent resumeLoopAgent stops on a pending call but cannot resume into the paused iteration; ParallelAgent has no resumable pathyes: loop index and count checkpointed, finished parallel branches skipped
EventActions.agentState / endOfAgentfields existwritten as checkpoints and rehydrated on resume
runAsync with an explicit invocationIdnoyes, @Experimental
plainTextContinuationAutoResumebuilder flag present@Deprecated legacy flow, mutually exclusive with resumable

The ADK documentation's resume page (adk.dev, runtime/resume), read on the same date, lists Python v1.16.0 and Kotlin v0.1.0 as supported and does not mention Java. Read that as documentation lagging the code, not as Java lacking the feature, and pin your ADK version deliberately.

Advertisement

Enabling it and building a pausable workflow

Resumability is configured once on the App and applies to every agent in it. The example is a three-step publishing flow: draft, ask a human for approval, publish. The approval step uses a long-running tool. The model names are placeholders; use whatever your project already runs.

import com.google.adk.agents.LlmAgent;
import com.google.adk.agents.SequentialAgent;
import com.google.adk.apps.App;
import com.google.adk.apps.ResumabilityConfig;
import com.google.adk.tools.Annotations.Schema;
import com.google.adk.tools.LongRunningFunctionTool;
import java.util.Map;

public final class PublishFlow {

  /** Starts a human approval and returns immediately; the answer arrives later. */
  public static Map<String, Object> requestApproval(
      @Schema(name = "draftId", description = "id of the draft to approve") String draftId) {
    String ticket = ApprovalQueue.open(draftId);       // your system; must be idempotent
    return Map.of("status", "pending", "ticket", ticket);
  }

  public static App build() {
    LlmAgent draft = LlmAgent.builder()
        .name("draft_agent").model("gemini-2.5-flash")
        .instruction("Write a release note draft and save it with an id.")
        .build();
    LlmAgent approval = LlmAgent.builder()
        .name("approval_agent").model("gemini-2.5-flash")
        .instruction("Ask for approval of the draft using requestApproval, then report the decision.")
        .tools(LongRunningFunctionTool.create(PublishFlow.class, "requestApproval"))
        .build();
    LlmAgent publish = LlmAgent.builder()
        .name("publish_agent").model("gemini-2.5-flash")
        .instruction("If the draft was approved, publish it; otherwise explain why not.")
        .build();

    SequentialAgent flow = SequentialAgent.builder()
        .name("publish_flow").subAgents(draft, approval, publish).build();

    return App.builder()
        .name("publish_app")
        .rootAgent(flow)
        .resumabilityConfig(ResumabilityConfig.builder().resumable(true).build())
        .build();
  }
}

Two things in that code are deliberate. The tool returns immediately with a pending status and a ticket; it does not block waiting for the human. And opening the ticket must itself be idempotent, for reasons the at-least-once section explains. Building with the release, the resumabilityConfig call produces a deprecation warning; suppress it locally with a comment explaining why, so the next reader knows it was a conscious choice.

Pausing and resuming on the released version

Run the flow. draft_agent produces its draft, approval_agent calls requestApproval, and the event carrying that call lists its id in longRunningToolIds. The resumable SequentialAgent sees the pending long-running call and stops; publish_agent does not run. Your code should record the pending call id next to the ticket, because that id is how the answer finds its way back.

import com.google.adk.artifacts.InMemoryArtifactService;
import com.google.adk.events.Event;
import com.google.adk.runner.Runner;
import com.google.genai.types.Content;
import com.google.genai.types.FunctionCall;
import com.google.genai.types.FunctionResponse;
import com.google.genai.types.Part;
import java.util.Map;

Runner runner = Runner.builder()
    .app(PublishFlow.build())
    .sessionService(persistentSessionService)   // NOT in-memory if you must survive restarts
    .artifactService(new InMemoryArtifactService())
    .build();

// Request 1: run until the long-running call pauses the sequence.
String pendingCallId = null;
for (Event e : runner.runAsync(userId, sessionId,
        Content.fromParts(Part.fromText("Prepare the 2.4 release note"))).blockingIterable()) {
  for (FunctionCall call : e.functionCalls()) {
    if (e.longRunningToolIds().map(ids -> ids.contains(call.id().orElse(""))).orElse(false)) {
      pendingCallId = call.id().orElseThrow();      // persist this with the ticket
    }
  }
}

// Request 2, possibly in another process days later: answer the paused call.
Part answer = Part.builder().functionResponse(
    FunctionResponse.builder()
        .id(pendingCallId)
        .name("requestApproval")
        .response(Map.of("status", "approved", "approver", "alice"))
        .build())
    .build();
runner.runAsync(userId, sessionId, Content.fromParts(answer))
    .blockingForEach(e -> log(e));                  // approval_agent resumes, then publish_agent

When the second request arrives, the runner looks back through the session for the function call that the new function response answers, and finds the agent that authored it. With resumability on, v1.10.1 routes the request to that agent's top-most SequentialAgent ancestor rather than straight to the agent, so the sequence can continue past it. The SequentialAgent then rebuilds its resume point from the session events: it finds which direct sub-agent's subtree made the call and starts from that index. draft_agent is not re-run; approval_agent sees the approval and reports it; publish_agent runs next.

Without resumability, the same function response would be routed straight to approval_agent, which would answer and stop, and publish_agent would never run. That difference is the clearest test that your configuration took effect (the sequential chain pattern covers the non-resumable baseline).

What main adds: durable checkpoints

Rebuilding the resume point from events works for a simple sequence, but it cannot capture state that never appears in an event, such as how many times a loop has run. On main, workflow agents write explicit checkpoints. A SequentialAgent writes an event whose agentState records current_sub_agent before each sub-agent runs; a LoopAgent records current_sub_agent and times_looped; a ParallelAgent checkpoints its start and skips branches already marked finished. An agent that completes is marked with endOfAgent, which drops its stored state.

On resume, the invocation context rehydrates those checkpoints from the invocation's events, each workflow agent fast-forwards to its checkpoint, and an invocation whose active agent already finished resolves to a no-op. Main also adds an experimental runAsync overload that takes the invocation id explicitly:

// Unreleased, main branch only, marked @Experimental (checked 2026-10-02).
// Resume a specific invocation; invocationId may be null when it can be inferred
// from a function response carried by the message.
runner.runAsync(
        userId,
        sessionId,
        invocationId,                  // the paused invocation, from Event.invocationId()
        Content.fromParts(answer),     // optional: may be null to just continue
        RunConfig.builder().build(),
        /* stateDelta= */ null)
    .blockingForEach(e -> log(e));
// Throws IllegalStateException if the App is not configured with resumable(true).

Two caveats come from the main-branch javadoc itself. Resume is best-effort and at-least-once. And in-memory state is lost on resumption: anything an agent keeps in a Java field rather than in session state or a checkpoint is gone after the pause. The older plainTextContinuationAutoResume flag, which made any plain user message resume the last unfinished invocation, is deprecated on main and cannot be combined with resumable; the builder rejects that combination.

At-least-once: why every tool must be idempotent

A crash can land between a tool doing its work and the event recording that work being persisted. On resume, the framework cannot know the work happened, so it runs the step again. The adk.dev resume page states the contract plainly: tools run at least once and may run more than once. For read-only tools that is harmless. For a tool that charges a card, sends an email or opens a ticket it is a duplicate side effect.

The fix is an idempotency key derived from business data, checked against a durable store before acting, and passed to downstream APIs that accept one:

public static Map<String, Object> chargeCard(
    @Schema(name = "orderId") String orderId,
    @Schema(name = "amountCents") long amountCents,
    ToolContext ctx) {
  // The key must be stable across retries and resumes: derive it from business data,
  // never from a random UUID generated inside the tool.
  String key = "charge:" + orderId;
  Optional<Receipt> done = receipts.find(key);       // durable store, not a field
  if (done.isPresent()) {
    return Map.of("status", "ok", "receipt", done.get().id(), "replayed", true);
  }
  Receipt r = payments.charge(orderId, amountCents, /* idempotencyKey= */ key);
  receipts.save(key, r);
  return Map.of("status", "ok", "receipt", r.id(), "replayed", false);
}

The key must be the same on every retry, which is why it comes from the order id and not from a UUID generated inside the tool. The receipt store must be durable and outside the agent process. Idempotency in ADK Java covers the pattern in more depth, including keys for tools whose arguments are generated by the model.

Operating resumable workflows

  • Use a persistent session service. The in-memory session service loses the event log when the process exits, and with it everything resumption needs. Persist events in a database; storing ADK events in Postgres shows one layout.
  • Persist the pending call id with your external ticket. The webhook or UI that receives the human's answer needs the session id, user id and function call id to build the function response.
  • Do not change the workflow while instances are paused. The adk.dev page lists this as a limitation. Renaming an agent or reordering sub-agents breaks the mapping from recorded authors and checkpoint names to code. Version workflows and drain paused instances before deploying structural changes.
  • Keep agent memory in session state. Values in Java fields do not survive the pause. Use state deltas, which are events and therefore durable.
  • Custom agents need work. Only the built-in workflow agents implement the resume logic. A custom BaseAgent subclass has to read and write its own checkpoints to resume correctly.
  • Expire stale pauses. Decide how long an approval may stay pending, and send a function response with a timeout or rejection status when it expires, so the workflow ends deliberately rather than never.

Failure modes

SymptomLikely causeFix
Function response is answered but the next agent never runsresumability not enabled, or the runner built from an agent instead of the Appbuild the Runner with app(...) and resumable(true)
Earlier agents run again after resumeworkflow agent without resume support, or a custom agentuse SequentialAgent on the release; add checkpoints to custom agents
Resume finds nothing to continuesession lost (in-memory service) or wrong session idpersistent session service; store ids with the ticket
Duplicate charges, emails or ticketsat-least-once re-execution after a crashidempotency keys and a durable receipt store
Resume errors after a deployagents renamed or reordered while pausedversion workflows; drain before structural changes
Builder throws on build() (main)resumable and plainTextContinuationAutoResume both setset only resumable

What to do next

  1. Pin your ADK Java version and read ResumabilityConfig and Runner in that exact version's source; the behaviour changes between releases.
  2. Build the three-agent flow above with a persistent session service and confirm that publish_agent runs after the function response, and does not without resumable(true).
  3. Kill the process between the two requests and resume from a fresh one to prove the pause really is durable.
  4. Audit every tool reachable from a resumable workflow for side effects, and add idempotency keys to each one that has them.
  5. Write down your policy for stale pauses and structural deploys before the first real paused instance exists.
Key takeaway: ADK Java resumability lets a workflow stop at a long-running tool call and continue in a later request from the agent that made it. In the v1.10.1 release this works for SequentialAgent by rebuilding the resume point from session events, under a config marked as a partial feature; main adds durable agentState checkpoints, loop and parallel support and an explicit invocation-id resume, all experimental. In every version the event log must be persistent, tools must be idempotent because resumption is at-least-once, and the workflow must not change shape while instances are paused.