Agents fail in ways RED metrics cannot see

The case for agent-specific observability starts with the failure taxonomy. Agents fail in ways that are invisible to RED metrics — rate, errors and duration all look perfect while the product is broken. A routing regression sends 30% of billing questions to the shipping specialist: every request returns 200, latency is normal, and users get confidently unhelpful answers. A prompt tweak makes a critic sub-agent too lenient and quality decays with zero errors. Compaction drops a commitment and the agent breaks a promise it made forty turns ago.

Each is detectable, but only by instruments aimed at behavior: sampled routing checks, escalation and handoff rates, guardrail-block counts. Teams that monitor an agent the way they monitor a web service learn about regressions from customers.

The debugging workflow inverts too. A stack trace localizes a crash; a wrong answer localizes nothing — the cause could be retrieval, context assembly, a misleading tool result, routing, or plain hallucination. The only tractable method is time travel over a complete record, which is why completeness is the load-bearing property of the stack.

Advertisement

Four lenses: events, traces, metrics, quality signals

A workable stack has four lenses, each answering a question the others cannot. The event log — ADK’s persisted Event stream per session, with authors, function calls and results, state deltas and control actions — is ground truth: what you replay and what you query for audit. Traces give structure per invocation: a root span for the turn, children for each model call and tool execution, so latency has a shape rather than a number. Metrics aggregate span attributes into the operational and economic view: latency percentiles, tool error rate, tokens and cost per session. Quality signals are the production proxies for ‘is it any good’ — escalation rate, guardrail blocks, thumbs-down rate, sampled human labels.

The mistake is treating these as four tools you buy. They are four views of one invocation, and the value is almost entirely in the joins. An architecture that cannot get from a cost anomaly to a trace to a replayable session has four dashboards and no answers.

Note the direction of dependency: traces, metrics and quality signals are all derived. The event log is the only one that is primary, and the only one whose loss is unrecoverable.

ADK observability — seeing inside a stochastic systemtraces for structure, events for truth, evals for qualityEvent logthe native audit recordOTel tracesspans per model/tool callToken metricscost per turn/agent/toolQuality signalseval scores, routing accuracyStructured loggingcorrelation ids everywhereDashboardslatency, cost, behaviorSession replaydebugging = time travelAlertingbehavioral + economicCallback instrumentationhooks as sensor pointsProduction miningfailed sessions → eval casesOps — PII-safe telemetry + sampling + trace/event correlationcorrelatevisualizereplaypageemitreviewimproveoperateoperate
ADK observability: the event log is ground truth, OTel spans give structure, token metrics give economics, and eval scores close the quality loop.
Advertisement

The correlation contract: one id set, everywhere

All of that depends on one unglamorous discipline: every telemetry artifact carries the same identifiers. ADK supplies the vocabulary — app_name, user_id, session_id, and the invocation_id scoping a single turn. Add a deployment_version and, in multi-tenant systems, a tenant id. That set is the contract: span attributes, log fields, and metric labels — minus the high-cardinality ones, because a metric labelled with session_id will bankrupt your time-series database.

IdScopeAnswers
trace_idone invocationshow this turn as a tree
invocation_idone turn (ADK)join spans, logs and events
session_idwhole conversationreplay it, sum its cost, audit it
user_id / tenantperson or accountattribute spend, honour deletion
agent name + versiona componentcompare canary against baseline

Enforce it in one place — a plugin and a log adapter — and assert it in CI. A contract individual developers must remember is a contract you do not have.

What a production span must carry

ADK’s runtime emits spans around the invocation, each agent run, each model call and each tool execution. The production question is narrower than how to read a waterfall: which attributes must be present for the on-call engineer and the finance rollup to work? Treat it as a schema.

SpanRequired attributesWhy
model callmodel id and version, input/output token counts, finish reason, cached tokens, durationcost, context bloat, truncation
tool calltool name, duration, status, error class, result sizetool error rate, oversized results that flood context
agent runagent name, version, correlation ids, step countper-agent SLOs, canary comparison
invocationend-to-end duration, time to first token, final statusthe latency the user actually felt

Two rules keep this honest. Record the error class — timeout, auth, rate_limit, not_found — not the error string, because free text is unaggregatable and often contains user data. And record result size alongside any content: size predicts context blowup and is safe to keep even when the payload is not.

The signals that actually page you

Most agent dashboards drown in charts nobody reads. A small set of SLIs carries almost all the operational signal, and each is derivable from the span attributes above.

SignalDefinitionWhy it moves
Time to first tokenmessage accepted to first streamed chunkperceived speed; regresses on prompt growth or a slow first tool
Turn latency p95invocation span durationreal completion time; hides step-count creep
Steps per turnmodel calls per invocationearliest sign of a loop or an over-decomposing planner
Tool error rateerrored tool spans / tool spans, by toola dependency degraded; the model is improvising
Tokens per sessionsum of model-span token countscost, context bloat and injection loops show here first
Guardrail block rateblocked / attempteda spike is an attack or a broken prompt

Note what is missing: request rate and HTTP error rate. Keep them, but they stay flat through every interesting agent incident. The two that pay for themselves fastest are steps per turn and tokens per session, because nearly every runaway failure inflates one of them long before a user complains.

Time to first token and the streaming budget

Turn latency is the honest number; time to first token is the one users feel. For a streaming agent the two diverge wildly: a nine-second turn that starts printing at 600 ms feels responsive, while a four-second turn showing nothing until 3.8 seconds feels broken. Measure only completion time and you optimize the wrong half.

Measuring TTFT in ADK means watching the event stream, not the return value. Iterating run_async yields partial events as the model streams, so the first event carrying text stops the clock.

t0 = time.perf_counter()
ttft = None
async for event in runner.run_async(
        user_id=uid, session_id=sid, new_message=msg):
    if ttft is None and event.content and event.content.parts:
        ttft = time.perf_counter() - t0
        span.set_attribute("agent.ttft_ms", ttft * 1000)
    if event.is_final_response():
        total = time.perf_counter() - t0

The gap between TTFT and total is your thinking budget. When it grows, the cause is almost always a tool call or planning step that now runs before the first token instead of after it — a design change, not a capacity problem, and visible only because the two numbers are tracked separately.

Structured logs that do not leak the user

Agent telemetry is unusually dangerous because the interesting payloads — prompts, tool arguments, tool results, model output — are exactly the fields containing personal data. ‘Log the request and response so we can debug it’ quietly turns your observability backend into an unmanaged copy of your user database, with a different retention policy and a wider access list.

The workable posture is an allowlist of shapes, not values: tool name, argument keys, result size, status and error class by default; values only for fields declared safe; hashes for identifiers you join on.

SAFE_ARG_KEYS = {"order_status", "region", "page"}

def tool_log_fields(tool_name, args, result, err=None):
    return {
        "tool": tool_name,
        "arg_keys": sorted(args),
        "args": {k: v for k, v in args.items()
                 if k in SAFE_ARG_KEYS},
        "result_bytes": len(json.dumps(result or {})),
        "error_class": type(err).__name__ if err else None,
    }

Full payloads still have a home — the session store, inside your trust boundary, under your retention policy — and debugging pulls them from there by session id.

Cost attribution: from token counts to a line item

A cloud bill tells you agents cost money. Attribution tells you which agent, which tool chain, which tenant, per what unit of value — a question only your own instrumentation answers, because the provider sees an API key, not your org chart. The raw material is the per-call token usage returned on every model response, and an after_model_callback — or the equivalent app-wide plugin hook — is the natural place to price it.

RATES = {  # USD per 1M tokens: (input, output)
    "gemini-2.0-flash": (0.10, 0.40),
}

def after_model(callback_context, llm_response):
    u = getattr(llm_response, "usage_metadata", None)
    if not u:
        return None
    inp = getattr(u, "prompt_token_count", 0) or 0
    out = getattr(u, "candidates_token_count", 0) or 0
    r_in, r_out = RATES.get(MODEL, (0.0, 0.0))
    metrics.record((inp * r_in + out * r_out) / 1e6,
                   agent=callback_context.agent_name,
                   invocation=callback_context.invocation_id)
    return None  # do not modify the response

Roll it up on four axes: per agent, per tool chain, per tenant, and per resolved task. That last ratio — cost per resolution, not cost per token — is the only one a business conversation can use, because a cheaper model that doubles turns is not cheaper.

Budgets, anomalies, and the runaway loop

Once cost is attributed per session, treat it as a saturation signal with a budget rather than a monthly surprise. Two controls do most of the work.

The first is a per-session budget guardrail. Accumulate spend in session state as it is incurred and have a before_model hook refuse to start another call once a ceiling is crossed, returning a graceful message or escalating to a human. Because ADK state is persisted with the session, the counter survives process restarts and spans the whole conversation, not one turn.

The second is anomaly detection on the distribution, not the mean. Mean token spend is nearly useless: a runaway loop in 1% of sessions can triple the bill while barely moving the average. Alert on p99 tokens per session, on steps per turn crossing a threshold, and on repeated identical tool calls within one invocation.

All three fire on the same pathology — the agent repeating itself because something in its input keeps saying ‘try again’. Catching it economically is usually catching it first, hours before the symptoms become complaints.

Plugins and callbacks as sensor points

You should almost never write instrumentation inside an agent. ADK offers two seams for cross-cutting telemetry, and choosing correctly decides whether coverage is a guarantee or a hope. Callbacks on an individual agent — before_model_callback, after_tool_callback and friends — suit instrumentation specific to that agent. Plugins registered on the runner fire the same hook points for every agent in the application, including agents added next quarter by someone who never read your logging guide. Policy — the correlation contract, cost accounting, redaction before export — belongs in a plugin.

HookEmit
before_modelassembled context size by component: instruction, history, tool results
after_modeltoken counts, cost, finish reason, model version
before_tooltool name, argument keys, budget and guardrail decisions
after_toolduration, result size, error class, sanitization hits

The before_model context-size histogram is the most underrated: it turns context engineering from guesswork into a chart showing which component eats the window.