An agent's tools are where it touches the world: databases, payment APIs, search, ticketing. When an agent is slow, wrong or expensive, the cause is usually a tool, and yet tool behaviour is the part teams instrument last. They trace the model call, see a turn take twelve seconds, and cannot say whether the time went to the model, to a slow downstream, or to a tool that failed three times and was retried by the model.
Recent ADK Java releases do more of this for you than most teams realise. The source of release v1.10.1 (published 18 September 2026) and of the main branch opens a span for every tool execution and records three histograms per call. This article shows exactly what is recorded, how to make sure it actually reaches your backend, what is missing, and how to add the outcome metrics and SLOs that turn raw telemetry into something you can page on.
What ADK records for every tool call
When the model returns a function call, the flow resolves the tool and wraps the whole execution, including the before and after callbacks, in a span named execute_tool <tool name>, for example execute_tool get_order. The tracer and the meter are both obtained from the global OpenTelemetry instance under the instrumentation name gcp.vertex.agent. The span carries GenAI semantic-convention attributes plus ADK-specific ones:
| Attribute | Value |
|---|---|
gen_ai.operation.name | execute_tool |
gen_ai.tool.name, gen_ai.tool.description | From the tool declaration |
gen_ai.tool.type | The tool's Java class simple name, for example FunctionTool |
gen_ai.tool_call.id | The id that ties the call to the model's function call |
gcp.vertex.agent.tool_call_args | Serialised arguments, if content capture is on |
gcp.vertex.agent.tool_response | Serialised response, if content capture is on |
gcp.vertex.agent.event_id | The function-response event id in the session |
When the model asks for several tools in one response, ADK subscribes to all of them eagerly (whether they actually overlap depends on whether each tool's work is asynchronous or on its own scheduler), each gets its own span, and ADK adds a span named execute_tool (merged) around the merged response event. Note that the README in the telemetry package still describes spans named tool_call [name] and tool_response [name]; the code is authoritative, so build dashboards on what your exporter actually receives. The ADK tracing walkthrough shows how these spans nest under the agent and model-call spans; it is written for the Python ADK, so its span names differ slightly.
When the span closes, ADK records three histograms, all with the attributes gen_ai.tool.name and gen_ai.agent.name:
| Instrument | Unit | What it measures | Extra attribute |
|---|---|---|---|
gen_ai.tool.execution.duration | ms | Wall time of the execution including callbacks | error.type = exception class simple name when the call threw |
gen_ai.tool.request.size | By | Size of the arguments | none |
gen_ai.tool.response.size | By | Size of the function response | none |
The same class records gen_ai.agent.invocation.duration and agent request and response sizes, so a single meter gives you agent-level and tool-level latency on the same axes. All of this was read from the release source; if you run an older version, check whether com.google.adk.telemetry.Metrics exists in your jar before relying on it.
The data flow, end to end
Two consequences follow from where the span sits. First, callback latency is charged to the tool: a slow policy check in a before callback shows up as a slow tool. Second, a before callback that returns a value short-circuits the tool, as the dispatch mechanics article explains, so a cache hit still produces a span and a duration sample, just a fast one. If you cache in a before callback, your p50 drops and your downstream's real latency becomes invisible unless you label hits separately.
Wiring OpenTelemetry so it is not a no-op
ADK depends only on the OpenTelemetry API. Without an SDK registered as the global instance, every span and histogram goes to a no-op implementation and nothing fails. That is the most common reason teams believe ADK emits no metrics. Order matters too. ADK obtains its meter and tracer from the global instance in static initialisers, and in OpenTelemetry Java the first read of an unset global installs a no-op global (unless SDK autoconfiguration is on the classpath and enabled). A later buildAndRegisterGlobal() then throws IllegalStateException saying set has already been called. Register the SDK before any ADK class loads, at the top of main.
import io.opentelemetry.exporter.otlp.metrics.OtlpGrpcMetricExporter;
import io.opentelemetry.exporter.otlp.trace.OtlpGrpcSpanExporter;
import io.opentelemetry.sdk.OpenTelemetrySdk;
import io.opentelemetry.sdk.metrics.Aggregation;
import io.opentelemetry.sdk.metrics.InstrumentSelector;
import io.opentelemetry.sdk.metrics.SdkMeterProvider;
import io.opentelemetry.sdk.metrics.View;
import io.opentelemetry.sdk.metrics.export.PeriodicMetricReader;
import io.opentelemetry.sdk.trace.SdkTracerProvider;
import io.opentelemetry.sdk.trace.export.BatchSpanProcessor;
import java.time.Duration;
import java.util.List;
public final class Telemetry {
public static void init() {
// Default SDK buckets stop at 10 s; tools that call other models or slow APIs need a longer tail.
View toolLatency = View.builder()
.setAggregation(Aggregation.explicitBucketHistogram(List.of(
5.0, 10.0, 25.0, 50.0, 100.0, 250.0, 500.0, 1000.0,
2500.0, 5000.0, 10000.0, 20000.0, 30000.0, 60000.0)))
.build();
SdkMeterProvider meters = SdkMeterProvider.builder()
.registerView(InstrumentSelector.builder()
.setName("gen_ai.tool.execution.duration").build(), toolLatency)
.registerMetricReader(PeriodicMetricReader.builder(
OtlpGrpcMetricExporter.getDefault()).setInterval(Duration.ofSeconds(30)).build())
.build();
SdkTracerProvider tracers = SdkTracerProvider.builder()
.addSpanProcessor(BatchSpanProcessor.builder(OtlpGrpcSpanExporter.getDefault()).build())
.build();
OpenTelemetrySdk.builder()
.setMeterProvider(meters)
.setTracerProvider(tracers)
.buildAndRegisterGlobal(); // must run before any ADK class touches telemetry
}
}The OpenTelemetry Java agent or SDK autoconfiguration achieves the same thing through environment variables; either is fine as long as the global instance exists first. Verify it the boring way: start the service, call one tool, and look for gen_ai.tool.execution.duration in your backend before writing any dashboard.
Content capture is on by default
ADK reads the environment variable ADK_CAPTURE_MESSAGE_CONTENT_IN_SPANS, and its default is true. With the default, tool arguments and tool responses are written into span attributes. For a get_order tool that means customer names, addresses and order contents end up in your trace store, which usually has wider access and longer retention than the database they came from. It also means large responses inflate span size, and many backends truncate or drop oversized attributes.
Set ADK_CAPTURE_MESSAGE_CONTENT_IN_SPANS=false in every environment that handles real user data, and turn it on deliberately in development. If you need arguments for debugging in production, log a redacted, allow-listed subset from a callback instead, so the decision about what leaves the process is explicit and reviewable.
What the built-in metrics miss
The built-in error.type is set only when the tool throws. Most well-behaved tools do not throw on business failures; they return a structured result such as {"status": "error", "retryable": false} so the model can explain it. Those calls record as successes. A tool that fails validation on 40 percent of calls can look perfectly healthy on the duration histogram.
The other gaps are outcome and cause: whether the call was a cache hit, a policy denial, a not-found, a downstream timeout, or a success. None of these are knowable to the framework. They are your contract, and the right place to record them is the tool callbacks, because they see every call to every tool without touching tool bodies.
Adding outcome metrics with callbacks
import com.google.adk.agents.LlmAgent;
import io.opentelemetry.api.GlobalOpenTelemetry;
import io.opentelemetry.api.common.AttributeKey;
import io.opentelemetry.api.common.Attributes;
import io.opentelemetry.api.metrics.LongCounter;
import io.reactivex.rxjava3.core.Maybe;
import java.util.Map;
import java.util.Optional;
final class ToolOutcomes {
private static final AttributeKey<String> TOOL = AttributeKey.stringKey("gen_ai.tool.name");
private static final AttributeKey<String> AGENT = AttributeKey.stringKey("gen_ai.agent.name");
private static final AttributeKey<String> OUTCOME = AttributeKey.stringKey("tool.outcome");
// Obtain after Telemetry.init() has registered the SDK.
private static final LongCounter CALLS = GlobalOpenTelemetry.getMeter("com.example.agents")
.counterBuilder("tool.calls").setDescription("Tool calls by outcome").build();
static String classify(Object response) {
if (response instanceof Map<?, ?> m) {
Object status = m.get("status");
if ("denied".equals(status)) return "denied";
if ("not_found".equals(status)) return "not_found";
if ("error".equals(status)) {
return Boolean.TRUE.equals(m.get("retryable")) ? "error_retryable" : "error_terminal";
}
}
return "ok"; // a small, closed set of values
}
static LlmAgent withOutcomes(LlmAgent.Builder builder) {
return builder
.afterToolCallbackSync((ctx, tool, args, toolCtx, response) -> {
CALLS.add(1, Attributes.of(TOOL, tool.name(),
AGENT, ctx.agent().name(), OUTCOME, classify(response)));
return Optional.empty(); // empty = keep the tool's own response
})
.onToolErrorCallback((ctx, tool, args, toolCtx, error) -> {
CALLS.add(1, Attributes.of(TOOL, tool.name(),
AGENT, ctx.agent().name(), OUTCOME, "exception"));
return Maybe.empty(); // empty = let the error propagate as before
})
.build();
}
}The callback interfaces used here come from the release source: the synchronous after callback receives the invocation context, the tool, the arguments, the tool context and the response, and returns an Optional map where empty means no override; the error callback receives the exception and returns a Maybe. If you short-circuit calls in a before callback, such as a cache or a policy gate, count those there with outcome cache_hit or denied, because the after callback sees the substituted value and cannot tell it apart from a real result unless you mark it. If your version lets an error callback turn an exception into a result, check whether the built-in duration still carries error.type for those calls before relying on it.
Worked example: one slow, flaky tool
The numbers here are illustrative, not measured. Suppose a support agent has three tools: get_order, search_kb and issue_refund. Users report that refunds sometimes take a minute. The traces show turns with three execute_tool issue_refund spans in a row. The duration histogram for issue_refund has a p50 of 400 ms and a p99 of 9.8 seconds, and the default buckets stop at 10 seconds, so the tail is flattened into the overflow bucket. After adding the view above, the p99 turns out to be 28 seconds.
The outcome counter shows 22 percent of issue_refund calls ending in error_retryable, with error.type empty because the tool never throws. The model is doing exactly what the tool told it: the error was retryable, so it retried, three times, inside one turn. The fix has two halves. On the downstream side, a per-tool timeout below the payment provider's own timeout turns a 28-second hang into a fast failure. On the contract side, the tool returns a terminal error after the first timeout for the same idempotency key, and the retry happens in the tool with backoff, where it is visible in metrics, instead of in the model, where it costs a full model round trip each time.
Cardinality, sampling and cost
- Keep metric attributes closed. Tool name, agent name, outcome class and error type are bounded. User ids, session ids, call ids, argument values and model-generated strings are not; each new value creates a new time series. Put those on spans, never on metrics.
- Watch for dynamic tool names. Toolsets that generate tool names from remote schemas, such as MCP servers, can add names at runtime. Alert on the count of distinct
gen_ai.tool.namevalues. - Sample traces, not metrics. Metrics are aggregated in process and are cheap at any call volume. Traces are not; use head sampling for normal traffic and keep all traces with errors if your pipeline supports tail sampling.
- Use the size histograms. A tool response is replayed on every later model call in the turn, so
gen_ai.tool.response.sizeis a leading indicator of token cost. A p95 over a few tens of kilobytes deserves a summarise-and-offload design.
SLOs and alerts
Pick one latency SLO and one success SLO per tool that has a user-visible effect. Exported to Prometheus, dots in names and attribute keys become underscores and the unit is appended to the name, so confirm the exact series names in your backend before copying these queries.
# p95 tool latency per tool over 5 minutes
histogram_quantile(0.95,
sum by (le, gen_ai_tool_name) (
rate(gen_ai_tool_execution_duration_milliseconds_bucket[5m])))
# share of calls that did not succeed, per tool (custom counter)
sum by (gen_ai_tool_name) (rate(tool_calls_total{tool_outcome!="ok"}[5m]))
/ sum by (gen_ai_tool_name) (rate(tool_calls_total[5m]))Alert on burn rate against the SLO, not on single spikes. Add one agent-level signal that tools alone cannot show: tool calls per turn. A sudden rise usually means the model is looping on a failing tool, and it shows up there before it shows up in cost. The ADK Java observability article covers the model-call side of the same dashboard.
Failure modes
- No SDK registered. Everything is a silent no-op. Test for the presence of the tool histogram in CI with an in-memory exporter.
- SDK registered too late. If ADK touched the global first, a no-op is already installed and registration fails at startup with
IllegalStateException. Initialise first thing inmain. - PII in traces. The content-capture default writes arguments and responses into spans. Turn it off for real data.
- Healthy-looking failures. Error maps are counted as successes by the built-in metric. Add the outcome counter.
- Flattened tails. Default buckets hide anything above 10 seconds. Register a view.
- Orphaned spans from your own async code. ADK propagates context across its RxJava operators; if your tool hops to its own executor, wrap tasks with
Context.current().wrap(...)or your child spans lose their parent.
What to do next
- Register the OpenTelemetry SDK globally at the top of main and confirm
gen_ai.tool.execution.durationreaches your backend after one tool call. - Set
ADK_CAPTURE_MESSAGE_CONTENT_IN_SPANS=falsein every environment with real user data. - Add a view with buckets up to at least 60 seconds for tool latency.
- Add an outcome counter from after and error callbacks, with a closed set of outcome values, and count before-callback short-circuits separately.
- Define a latency and a success SLO for every tool with a user-visible side effect, and alert on burn rate.
- Chart tool calls per turn and response size per tool, and review the top three tools by each monthly.