Trace capture and replay
Evaluating agents manually on every model update does not scale. Instead, capture every execution trace — every prompt, response, tool call, and latency — and replay those same traces against new model versions, prompt templates, and system configurations. This becomes your regression test suite. A change that improves latency but breaks accuracy shows up immediately; a model update that silently degrades edge-case handling surfaces before users see it.
The capture phase is straightforward: instrument your agent to log JSON or structured traces at every step. Record the initial request, the exact prompt text sent to the LLM, the model version and temperature, the response, any tool calls made, their results, and the final output. Crucially, capture non-deterministic behavior: if your agent uses randomization or sampling, log the random seed or the exact sampled value. This lets you replay that execution identically even though the underlying system might behave differently.
Replay is where scale multiplies value. Given a trace, you can run the agent again using only the captured data as input, bypassing the LLM and tools entirely for fast, cost-free regression. Did the final output change? Was the reasoning path different? Did latency improve? A baseline run of 10,000 traces takes minutes instead of hours and costs nearly nothing. Tools like Langfuse, Arize Phoenix, and OpenLLMetry all provide this capability baked in.
The catch is maintaining trace fidelity. If you modify tool signatures or change how the agent interprets responses, old traces may no longer be valid. Versioning traces and keeping them alongside code changes (in the same commit, same branch) prevents silent mismatches. A trace replay suite that runs in CI catches regressions before they ship; traces that sit in a separate system and nobody checks are just storage.
LLM-judge with calibration
Human evaluation does not scale past hundreds of traces. Instead, delegate grading to a strong LLM judge — typically GPT-4, Claude 3 Opus, or Llama 3.1 405B — using a structured rubric. The judge reads the original request, the agent’s final output, and any intermediate steps, then scores the response on dimensions you care about: correctness, completeness, tone, safety, hallucination, tool-use accuracy.
The critical step is calibration. An LLM judge is not truth; it is a proxy. Early in your evaluation pipeline, manually grade a representative sample of 100–500 traces yourself or with a small team. Then run the judge on the same sample and compute the correlation between judge scores and human scores — Pearson correlation for continuous scores, Cohen’s kappa or accuracy for categorical judgments. Target >0.7 for Pearson correlation; anything lower means the judge is drifting from your intent.
Calibration is not a one-time affair. When you switch to a newer model (GPT-4 to GPT-4.5, Claude 3 Sonnet to Claude 3.5 Sonnet), re-run the calibration on your golden set because the judge’s biases and failure modes may have shifted. Trends in judge accuracy across your test set over time are an early warning that something in the judge’s environment (context window, system prompt tuning, or fine-tuning of the underlying model) has drifted. Document your calibration data as part of your evaluation setup so that later team members understand why specific rubric phrasings exist.
The rubric itself is the bulwark against judge unreliability. Write rubrics as if you were training a human grader: include concrete examples of high, medium, and low scores, anti-patterns that trip up novices, and edge cases you care about. A rubric like “is the output good?” will produce noisy scores; a rubric like “does the output answer all three sub-questions asked, use only facts from the provided documents, and avoid hedging language when a definitive answer exists?” grounds the judge in observable, testable criteria.
Drift monitoring
Evaluations in development are not enough. In production, track quality scores over time on every request, sampled or in full. Plot the rolling average and percentiles: if median quality drops from 0.82 to 0.74 overnight, something changed. Your job is to quickly distinguish between two very different causes and respond accordingly.
Sudden drops (hours to a day) usually mean the model provider changed something: a silent retraining run, a rollout of a new version, a shift in sampling behavior. Check the provider’s status page, ask in community channels, or roll back to an earlier model checkpoint if available. This has happened with Claude, GPT, and Llama updates; it is not rare. Sudden shifts also appear after you deploy a code change (a new prompt, a tool signature change, a system message tweak) — which is why production evaluation reveals bugs that staging does not.
Slow degradation (days to weeks) signals input distribution shift. Your evaluation dataset or golden test set was collected at a point in time; production traffic may be drifting. Users may be asking different kinds of questions, tools may be returning different data, or domain conventions may be evolving. A slow drop is harder to act on but more common. The response is to continually collect new traces from production, hand-label a sample, and re-calibrate your judge or refresh your baselines.
Set up automated alerts on your quality percentiles. A p50 drop >5% over a rolling 24-hour window, or a p95 drop >10%, typically warrants a page. Without alerts, you find out about drift from users, and by then the damage is done. Couple alerts with dashboards that let on-call engineers quickly isolate which input types, tool calls, or reasoning paths broke, so they can either rollback a change or escalate to the eval team for investigation.
Building reliable eval datasets
No evaluation strategy is better than its dataset. A small set of hand-picked examples is biased; a large set of low-quality labels is noisy. The goal is a diverse, representative, and correctly-labeled dataset that you trust to predict production behavior.
Start with stratified sampling. Identify dimensions your agent cares about: request complexity (simple vs. multi-step), domain area (if your agent handles multiple domains), language (if multilingual), edge cases (requests that trip up naive systems). Sample uniformly or stratified across these dimensions so your eval set does not over-represent easy cases or one domain. If your agent fails spectacularly on requests involving dates or arithmetic, make sure those are well-represented in your evals.
Hand-label a core golden set of 200–500 examples yourself or with domain experts. For each request, provide the correct answer or acceptable answer range. If an answer can be subjective (tone, completeness), define the criteria clearly in advance and discuss edge cases with anyone else labeling, so you reach high inter-rater agreement. Once you have 200 golden examples labeled, run your judge on them and check correlation; if it is <0.6, your rubric is ambiguous or your examples are too hard.
Once the golden set is solid, you can use it to bootstrap larger evals via weak supervision or semi-supervised learning, or simply run your judge on a larger unlabeled set. But always keep the golden set clean and human-verified; it is your reference point and your canary. If the judge drifts away from the golden set, you catch it immediately.