A SequentialAgent runs its sub-agents one after another. That is ideal for pipelines, until a pipeline touches the outside world. Suppose stage two reserved stock, stage three charged a card, and stage four, shipping, fails. The run ends with an error, the customer has been charged for an order that will never ship, and nothing in ADK will undo it.

This article explains why that happens, why the session cannot simply be rolled back, and how to build compensation properly: side effects in deterministic stages, a compensation record for each one, a small custom agent that undoes completed steps in reverse order, idempotency keys that make retries safe, and a durable ledger for the crash that happens at the worst moment. The same ideas are known as the saga pattern in distributed systems; here they are applied to agent pipelines.

Advertisement

What a SequentialAgent does when a stage fails

A sequential agent concatenates the event streams of its stages: stage two starts when stage one's stream completes. Each event can carry a state delta, and the runner appends every event to the session as it is emitted. If any stage signals an error, for example because a tool threw or an HTTP call timed out in a custom stage, the combined stream terminates with that error. Later stages never start.

Two things do not happen. The events that were already appended stay in the session, so charge_id and everything else written by earlier stages remains in state. And nobody calls the payment provider back. The stage design in sequential agent chains shows how to halt a chain on a failed check; halting stops future damage but does nothing about damage already done.

Also consider soft failures. An LlmAgent that decides the order is invalid does not throw; it writes text. If that verdict should trigger a rollback, you need a deterministic way to turn it into a failure, which is what the saga_abort key below is for.

Why there is no rollback

Rollback works in a database because the database owns every change and can keep undo information until commit. An agent pipeline owns nothing it changes. The card network, the warehouse system and the email provider each commit independently the moment you call them. There is no transaction to abort.

The session does not help either. It is an append-only event log, and state is the result of replaying deltas. You could write deltas that reset keys, but that only rewrites your notes about the world, not the world. The only way to undo an external effect is to perform another external effect with the opposite meaning: release the reservation, refund the charge, cancel the shipment. These are compensations, and they have properties a rollback does not:

  • They are new actions that can themselves fail, time out or be retried.
  • They are often semantic rather than exact: a refund is not an un-charge, since the customer sees both lines on the statement.
  • Some actions have no compensation at all. A sent email cannot be unsent; you can only send a correction.
  • They run later, so the world may have moved: the shipment may already be at the depot.
Advertisement

The architecture: effects in code, compensations recorded

The design has four parts. First, LLM stages decide and deterministic stages act: models plan the order and draft messages, while stages written in Java call the payment and inventory APIs. That keeps every side effect in code you can reason about, and it means the compensation record is written by code, not by a model that might forget. Second, each acting stage emits, in the same event as its result, a comp: entry describing how to undo it. Third, a custom agent, SagaSequence, runs the stages, collects those entries, and on failure calls the matching compensator for each, newest first. Fourth, a durable ledger outside the session covers the case where the process dies mid-step.

A SequentialAgent run that fails at stage 4, and the compensations that undo stages 3 and 21. planLlmAgent, no effect2. reservestock held3. chargecard charged4. shipcarrier refuses5. notifynever runsreleaseundo reserverefundundo chargeSagaSequencecatches the errorerrorCompensations run newest first, each with its own idempotency keySession event log (append-only)forward deltas, then comp records, then saga_status = compensatedDurable intent ledgersurvives a crash between the side effect and the event
Forward stages record how to undo themselves. When stage 4 fails, SagaSequence catches the error and runs the recorded compensations newest first, then records the outcome in the session.

An acting stage that records its own undo

This stage uses the same BaseAgent and Event building blocks as the sibling articles. It reads the plan an earlier LlmAgent wrote with outputKey (see passing context between steps), charges the card with an idempotency key derived from the invocation, and emits the charge id together with its compensation record:

import com.google.adk.agents.BaseAgent;
import com.google.adk.agents.InvocationContext;
import com.google.adk.events.Event;
import com.google.adk.events.EventActions;
import io.reactivex.rxjava3.core.Flowable;
import io.reactivex.rxjava3.schedulers.Schedulers;
import java.util.List;
import java.util.Map;
import java.util.concurrent.ConcurrentHashMap;

/** Deterministic stage: charges the card, then records how to undo it. */
public final class ChargeStage extends BaseAgent {
  private final Payments payments;          // your client; charge() takes an idempotency key

  public ChargeStage(Payments payments) {
    super("charge", "Charges the customer for the planned order.", List.of(), List.of(), List.of());
    this.payments = payments;
  }

  @Override
  protected Flowable<Event> runAsyncImpl(InvocationContext ctx) {
    return Flowable.fromCallable(() -> {
      Order order = Order.parse(ctx.session().state().get("order_plan"));
      String key = ctx.invocationId() + ":charge";            // same key on every retry
      String chargeId = payments.charge(order.customerId(), order.total(), key);

      Map<String, Object> delta = new ConcurrentHashMap<>();
      delta.put("charge_id", chargeId);
      delta.put("comp:charge", Map.of("step", "charge", "charge_id", chargeId,
                                      "key", ctx.invocationId() + ":refund"));
      return Event.builder()
          .id(Event.generateEventId())
          .invocationId(ctx.invocationId())
          .author(name())
          .branch(ctx.branch().orElse(null))
          .actions(EventActions.builder().stateDelta(delta).build())
          .timestamp(System.currentTimeMillis())
          .build();
    }).subscribeOn(Schedulers.io());                           // blocking HTTP off the Rx thread
  }

  @Override
  protected Flowable<Event> runLiveImpl(InvocationContext ctx) {
    return Flowable.error(new UnsupportedOperationException("live mode not supported"));
  }
}

The idempotency key matters twice. If the charge call times out and is retried, the provider returns the original charge instead of creating a second one. And the refund has its own key, so a compensation retried after a crash cannot refund twice. Idempotency for agent tool calls covers key design in depth. The compensation record is emitted only after the charge succeeded; the gap that leaves is the subject of the durable ledger section.

The SagaSequence agent

The orchestrator runs its stages with concatMap, exactly as a sequential agent does, and watches every event's state delta. Compensation records go on a stack; a saga_abort key turns a soft failure into an error. When the stream fails, it pops the stack and calls each compensator:

/** Runs stages in order. On failure, runs recorded compensations newest first. */
public final class SagaSequence extends BaseAgent {
  static final String COMP_PREFIX = "comp:";
  static final String ABORT_KEY = "saga_abort";       // a stage may request a rollback

  private final Map<String, Compensator> compensators;  // step name -> undo action

  public SagaSequence(String name, List<? extends BaseAgent> stages,
                      Map<String, Compensator> compensators) {
    super(name, "Sequential stages with compensation on failure.", stages, List.of(), List.of());
    this.compensators = Map.copyOf(compensators);
  }

  @Override
  protected Flowable<Event> runAsyncImpl(InvocationContext ctx) {
    return Flowable.defer(() -> {
      Deque<Map<String, Object>> done = new ArrayDeque<>();     // per run, newest on top
      return Flowable.fromIterable(subAgents())
          .concatMap(stage -> stage.runAsync(ctx).doOnNext(e -> {
            Map<String, Object> d = e.actions().stateDelta();
            d.forEach((k, v) -> {
              if (k.startsWith(COMP_PREFIX) && v instanceof Map<?, ?> m) done.push(copy(m));
            });
            if (d.containsKey(ABORT_KEY)) throw new SagaAbort(String.valueOf(d.get(ABORT_KEY)));
          }))
          .concatWith(Flowable.defer(() -> Flowable.just(status(ctx, "committed", null))))
          .onErrorResumeNext(err -> Flowable.fromCallable(() -> compensate(ctx, done, err))
              .subscribeOn(Schedulers.io()));
    });
  }

  private Event compensate(InvocationContext ctx, Deque<Map<String, Object>> done, Throwable cause) {
    List<String> failed = new ArrayList<>();
    while (!done.isEmpty()) {
      Map<String, Object> rec = done.pop();
      String step = String.valueOf(rec.get("step"));
      try {
        compensators.get(step).undo(rec);                // idempotent, retried inside
      } catch (Exception ex) {
        failed.add(step);                                // keep going; report at the end
        DeadLetters.park(ctx.invocationId(), rec, ex);
      }
    }
    return status(ctx, failed.isEmpty() ? "compensated" : "compensation_failed",
                  cause.getClass().getSimpleName() + (failed.isEmpty() ? "" : " " + failed));
  }
  // status(...) builds an Event whose stateDelta holds saga_status and saga_reason,
  // exactly like ChargeStage builds its event. runLiveImpl returns Flowable.error(...).
}

Several choices here are deliberate. The stack is created inside Flowable.defer, so each run gets its own and concurrent sessions never share one. The agent keeps compensating after one compensation fails, because leaving the refund undone just because the release failed is worse. Failed compensations go to a dead-letter store for a human or a sweeper, as in dead-letter queues for agent pipelines. And the error is converted into a final event with saga_status rather than rethrown, so the caller sees a clean, recorded outcome. Rethrow instead if your caller must treat a compensated run as an error.

Wiring it up replaces SequentialAgent.builder() for this pipeline:

static final SagaSequence ORDER_SAGA = new SagaSequence("order_saga",
    List.of(PLAN,                              // LlmAgent, outputKey("order_plan"), no side effects
            new ReserveStage(inventory),       // writes comp:reserve
            new ChargeStage(payments),         // writes comp:charge
            new ShipStage(carrier),            // writes comp:ship, or throws
            NOTIFY),                           // LlmAgent drafting the customer email, last on purpose
    Map.of("reserve", rec -> inventory.release(str(rec, "reservation_id"), str(rec, "key")),
           "charge",  rec -> payments.refund(str(rec, "charge_id"), str(rec, "key")),
           "ship",    rec -> carrier.cancel(str(rec, "shipment_id"), str(rec, "key"))));

Worked example: an order that cannot ship

StepEvent and stateWorld
planorder_plan = SKU-17 x 2, total 58.00No change
reservereservation_id = R-881, comp:reserve pushed2 units held
chargecharge_id = CH-42, comp:charge pushedCard charged 58.00
shipcarrier call fails: address not serviceableNo shipment
compensate 1refund CH-42 with key inv:refund58.00 refunded
compensate 2release R-881 with key inv:releaseStock back on shelf
finalsaga_status = compensated, saga_reason = ShipFailureConsistent

The notify stage never runs, because an LlmAgent placed after the failing stage is never reached. That placement is intentional: the email is the one action that cannot be compensated, so it goes last, after every compensatable step has succeeded. This is the pivot rule: order stages so that everything before the point of no return can be undone, and everything after it can only be retried until it succeeds. If the customer must be told about the failure, add a separate apology step that runs on saga_status = compensated.

Stages in a ParallelAgent nested inside the saga work the same way, because the orchestrator sees events from every branch. Their compensations can run in any order relative to each other; only order relative to earlier and later stages matters. How a branch failure surfaces is covered in parallel agents handling failures.

The crash window and the durable ledger

The in-memory stack has a hole. If the process dies after the payment provider accepted the charge but before the event reached the session, no compensation record exists anywhere, and the customer is charged for nothing. Restarting the pipeline does not help, since a new invocation has a new id and new keys.

Close the hole by writing intent before acting. Before calling the provider, the stage inserts a row into a durable table: invocation id, step, idempotency key, status PENDING. After the call it marks the row DONE with the charge id, and a final saga status marks the whole invocation COMMITTED or COMPENSATED. A sweeper job looks for invocations stuck with PENDING or DONE rows older than a timeout, asks the provider what happened using the idempotency key, and compensates. The session is then a convenient view, and the ledger is the source of truth. If the stage also publishes events, write the row and the outbox record in one transaction, as in the transactional outbox.

Timeouts create the same ambiguity without a crash: a charge that timed out may have succeeded. Compensators must therefore handle unknown state: look up the charge by key, refund it if it exists, and treat not-found as already compensated.

Failure modes

SymptomCauseFix
Charged but no shipment, no refundPlain SequentialAgent, no compensationSagaSequence with recorded compensations
Double refundCompensation retried without a keyDerive the refund key from the invocation
Orphan charge after a restartCrash between effect and eventIntent ledger plus sweeper
Email sent for a cancelled orderUncompensatable step before the pivotMove it last
Model decides abort, nothing undoneSoft failure never becomes an errorDeterministic check writes saga_abort
Compensation fails foreverResource already gone or changedDead-letter it, alert, human decision
Two runs share a stackState held in an agent fieldCreate per-run state inside defer

Trade-offs and operating it

Compensation is not free. Every acting stage needs an undo, a key scheme and tests for both, and the customer can see the intermediate state: a pending charge, then a refund. If an operation can be deferred instead, prefer that: reserve with an expiry and capture payment only at the end, so failure means doing nothing. Use compensation where a step must commit before the next can start.

Keep LLM calls out of compensators. Undo must be deterministic and fast, and a model deciding how to refund is a new failure mode during the worst possible moment.

Operationally, emit a metric for every saga outcome, committed, compensated and compensation_failed, and alert on the last one immediately. Track time from failure to fully compensated. Test with fault injection: make each stage fail in turn, including after its side effect but before its event, and assert that the world ends consistent. Run the sweeper in staging with a killed process at least once before trusting it.

What to do next

  1. List every external side effect your sequential pipeline performs and write its compensation, or mark it as uncompensatable.
  2. Move uncompensatable steps after the pivot, as the last stages.
  3. Move side effects out of LLM tool calls and into deterministic stages that emit comp: records.
  4. Derive idempotency keys from the invocation id for both actions and compensations.
  5. Replace the SequentialAgent with a SagaSequence and add a saga_abort check for soft failures.
  6. Add an intent ledger and a sweeper for the crash window.
  7. Inject a failure at every stage in tests, and alert on compensation_failed.
Key takeaway: A SequentialAgent stops at the first failing stage and leaves every completed side effect in place. Neither ADK nor the append-only session can roll back the outside world, so you compensate: deterministic stages act with idempotency keys and emit a record of how to undo themselves, and a small SagaSequence agent runs those undos newest first when a later stage fails or a check writes saga_abort. Put uncompensatable steps last, back the session with a durable intent ledger and a sweeper for crashes, dead-letter compensations that fail, and test by injecting a failure at every stage.