Product choice, not architecture
Building the actual tracing architecture -- spans per LLM call and per tool call, parent-child relationships across a plan-execute-reflect loop -- is covered by this series' observability and tracing architecture article. This one is narrower: given that architecture, which product should capture and store the resulting traces? The honest answer changes depending on whether a team already has observability infrastructure worth extending, or is starting from nothing.
Purpose-built LLM observability platforms
Products in this category (LangSmith and Langfuse are the reference examples; Arize represents the ML-observability-first end of the same space) are built around the specific shape of LLM traffic: a trace is a tree of prompt/response pairs, each carrying token counts, cost, latency, and often the full message history that produced a given completion. They ship with UIs purpose-built for reading that shape -- diffing two runs of the same prompt, filtering traces by cost or latency outliers, and replaying a specific failing trace with a modified prompt to see if a fix would have worked.
That last capability -- prompt-level replay against a captured trace -- is the feature that's genuinely hard to replicate with generic infrastructure, because it requires the tool to understand that a span is specifically an LLM call with a prompt template and variables, not just an opaque unit of work. If prompt iteration against real production failures is a frequent activity for the team, this category earns its keep specifically for that workflow.
The cost is a new vendor dependency and, for the hosted versions of these products, another system with access to potentially sensitive prompt/response content -- the same data-locality consideration that applies to any managed service handling application data.
Building on OpenTelemetry
The alternative is treating an LLM call as just another span in a standard OpenTelemetry trace, using the same collector, backend, and dashboards already handling the rest of the application's observability. Concretely: wrap each LLM call and tool call in a span, attach prompt, response, token count, and cost as span attributes, and let the existing trace-visualization tooling show the resulting parent-child tree alongside spans for database calls, HTTP requests, and everything else already instrumented.
What this buys: one observability stack instead of two, and traces that show an LLM call in the same timeline as the database query or downstream API call that surrounds it -- useful when the actual latency culprit in a slow agent run turns out to be a tool call, not the model, and a unified trace makes that visible at a glance instead of requiring a correlation across two separate systems.
What it costs: the prompt-specific UX -- diffing prompt variants, replaying a trace with a modified prompt -- has to be built by hand if it's wanted at all, and most generic tracing backends have no native concept of "this span's cost was $0.003" without custom dashboarding on top of the raw attribute.
What to actually capture, regardless of which product you pick
The tool matters less than the discipline of capturing the same fields consistently, because a trace missing one of these fields is a trace that can't answer the question it was built for.
per LLM-call span:
prompt (template id + resolved variables, not just final string)
response (raw, before any post-processing)
model + provider, token counts (input/output), cost, latency
parent span id (which plan step / tool call triggered this)
per tool-call span:
tool name, arguments, result, success/failure, latency
retry count if a retry policy fired
per run (root span):
goal / user request, final outcome, total cost, total latency,
replan count if applicableThe prompt template id plus resolved variables (not just the final rendered string) is the detail most setups skip and most regret skipping -- it's what makes "show me every trace that used prompt version 14" a query instead of a manual grep, and it's the exact field a purpose-built platform's replay feature depends on. Capture it from day one, on either kind of tooling, and the choice between the two stops being a decision you can't walk back.