Locate Failed Calls: First Fault, OpenTelemetry Beats Agent Logs

TakeawayDetail
Tracing should own first-fault localization.The OTelBench leader passed only 29% of instrumentation tasks, and context propagation was the most prominent failure mode. Complete parent-child topology locates the first causal edge needing inspection.
The 80.9% contrast cannot rank traces over logs.The same model scored 80.9% on SWE-Bench, but the benchmarks measured different work and did not compare traces with agent logs. Logs still explain payload and validator mismatches.
Observability complexity is a major operational obstacle.Quesma says 39% of organizations cite observability complexity as their top obstacle, reinforcing the need for a coherent causal map before detailed analysis.
Failed-call localization has material economic stakes.Quesma cites an average enterprise-outage cost of $1.4 million per hour, making accurate first-fault isolation an operational priority.

Quesma’s OTelBench release reports a striking split: the leading model passed just 29% of OpenTelemetry instrumentation tasks, while the same model posted an 80.9% pass rate on SWE-Bench. The tasks differ, so the scores are not apples-to-apples. Still, the result exposes a hard problem: a model can produce capable code while failing to preserve the execution context that makes distributed calls investigable.

Open on a successful HTTP tool response. The transport succeeded, but an agent validator can still reject the result because the arguments, schema, state transition, or business intent is wrong. A complete parent-child span topology shows where the call sits in the causal chain and identifies the first edge needing inspection. For locating that fault, OpenTelemetry is the default winner; logs explain what a technically successful but semantically wrong call contained.

The operating rule is simple: trace to the first causal edge, then use targeted logs for payload and validation evidence at that edge. Quesma identifies context propagation—the parent-child relationship between services—as the benchmark’s most prominent failure mode. That supports the mechanism, not a universal traces-over-logs ranking. Quesma also reports that 39% of organizations cite observability complexity as their top obstacle. Complete topology narrows the search before semantic evidence is read.

Locate Failed Calls

First Failing Edge

First isolate the earliest child on a root-to-leaf causal route that errors or times out—not the earliest log line or the last attempt that succeeds. Represent each agent turn as one OpenTelemetry trace: an agent root span owns children for every model generation, retrieval call, MCP client/server exchange, and external tool invocation. Pin the deployed OpenTelemetry GenAI and MCP semantic-convention versions in the instrumentation manifest; otherwise, a backend upgrade can change the fields used by the isolation query.

Propagation, not span creation, is the decisive weak link. According to the official OpenTelemetry documentation, context propagation is a separate core concept: emitting spans does not preserve continuity unless every hop accepts and forwards W3C Trace Context. According to Quesma’s January 2026 OTelBench press release, inter-service propagation was the benchmark’s most prominent failure mode and an “insurmountable barrier” for most tested models. Carry the 16-byte trace ID and 8-byte span ID across every model, retrieval, MCP, and tool boundary, then copy both into the OpenTelemetry Log Data Model record. A body event can thereby identify one invocation rather than merely the overall turn.

Make every retry a distinct child span, with its own start time, duration, status, attempt number, and link to the logical operation. If a retriever times out and a later attempt succeeds, the terminal success must not overwrite the failed attempt or hide its downstream effect on generation. Stable call and attempt identities let the analyst compare attempts and follow the failing attempt’s descendants.

For a technical invocation failure, set the span status to ERROR, attach an exception event, and make that span canonical. A long free-text log cannot reconstruct causality automatically; without W3C context and stable call or attempt IDs, it is only a pile of correlated symptoms. Task-policy validation is different: if transport completed but the prompt, retrieval, tool arguments, or output is semantically wrong, emit a separate redacted structured agent-log event and make it canonical, linked by trace_id and span_id. That separation prevents transport completion from being mistaken for task completion.

Isolation is a dependency-graph operation. Start at the agent root, follow parent-to-child causal edges, select the earliest child with ERROR or timeout status on each relevant route, and inspect its downstream effects. Do not sort emitted events by wall-clock timestamp: concurrent work makes clock order a poor proxy for causal order. This structure makes the benchmark concrete—an analyst opens the root and failing-edge view, then expands that edge and its descendants—to test first-fault localization in no more than two analyst actions on fully instrumented transport or dependency failures.

Symptom at the first edge Canonical record Required linkage Isolation action
Error, timeout, or bad dependency status OpenTelemetry ERROR span plus exception event 16-byte trace ID, 8-byte span ID, and parent edge Select the earliest causal failure, then inspect descendants
Retry followed by technical success or failure One distinct OpenTelemetry child span per attempt Attempt number and logical-operation link Inspect every attempt rather than accepting the terminal state
Transport-clean but wrong prompt, retrieval, arguments, or output Redacted structured agent-log event trace_id and span_id for the successful invocation Review the exact linked edge and its task-policy evidence
First Failing Edge — Locate Failed Calls

ATR-26

ATR-26 turns the article’s thesis into a falsifiable replay experiment rather than an instrumentation anecdote. Its recorded failure episodes form a full factorial: DNS, connection, TLS, HTTP rate limiting, HTTP 5xx, and downstream timeout, each occurring at the model API, retriever, MCP server, or downstream tool, with 25 deterministic seeds per fault–call pairing. Identical requests and fault schedules are replayed through trace-only, structured-log-only, and trace-indexed-log views. The primary target remains the earliest erroneous invocation on the causal route; transport-clean, semantically wrong outputs are outside this technical-failure denominator.

According to Qin et al.’s ToolBench (ICLR 2024), the tool-diversity baseline provides an API inventory. ATR-26 samples schemas from that inventory, but treats ToolBench strictly as a workload source—not as evidence that either telemetry view wins. According to Liu et al.’s AgentBench (ICLR 2024), the orchestration baseline evaluates eight agent environments. First-fault accuracy must therefore be reported separately for every environment; a pooled result cannot allow one domain to decide the article.

The reproducibility manifest must pin one exact OpenTelemetry Collector release and immutable image digest before the first run. The supplied protocol identifies no release number, so none should be inferred or fabricated. Its tail-sampling processor reference settings include decision_wait=30s. These are pipeline parameters, not measured agent overhead; their runtime effect is assessed only through the measured p95 overhead endpoint.

Each arm receives the same recorded request and is scored on first-fault accuracy, seconds to isolate, evidence bytes inspected, retained-error rate, and p95 execution overhead. An analyst action is exactly one query, filter, span expansion, or log click. Longer free-text logs do not become causal evidence automatically: without W3C trace context and stable call/attempt IDs, they remain piles of correlated symptoms. The trace-indexed-log arm may attach a redacted event to a span through trace_id and span_id, but that enrichment cannot rescue the log-only arm.

Preregistered endpointPass conditionDecision
First-fault resolutionThe preregistered first-fault resolution criterion is met in no more than two analyst actionsTrace-only must pass
Accuracy advantageAt least 10 percentage points above structured-log-only first-fault accuracyOpenTelemetry is the required technical winner
Execution costP95 overhead remains within the preregistered limit against an uninstrumented replay baselineTrace instrumentation must remain bounded
Supporting measurementsPublish observed seconds, evidence bytes, retained-error rate, and execution overhead with confidence intervalsNo endpoint may be omitted after collection
VerdictAll three gates passFailure of any gate falsifies the stated thesis

Freeze the Collector manifest, ToolBench schema sample, AgentBench environment strata, deterministic seeds, and scoring code before replay. Publish the observed values and confidence intervals regardless of outcome; favorable point estimates cannot erase a failed preregistered threshold.

Locate Failed Calls, photo 2

First-Fault Scorecard

OpenTelemetry wins the call-isolation contest; structured agent logs become canonical only when execution is transport-clean and the task is semantically wrong. The dividing line is causal reconstruction versus evaluative evidence: spans show which invocation belongs under which parent and how attempts relate, while logs can explain what a technically successful invocation was supposed to mean.

Diagnostic question OpenTelemetry trace Structured agent log Winner
Where did the call fail in the agent path? Explicit parent/child span tree and links Events must be reconstructed through IDs and timestamps OpenTelemetry
Which retry first failed? Separate attempt spans linked to one logical operation Retry events may be absent, duplicated, or interleaved OpenTelemetry
Why was an otherwise successful call wrong? Sequence is visible, but semantic correctness is not intrinsic Validator event can expose the bad argument, retrieval, or policy interpretation Agent log
Where is the full prompt or tool body? Storing it in attributes creates excessive span size Structured body or secure reference is natural Agent log
Overall task: isolate the first failed invocation Direct causal topology plus failure status Post-hoc reconstruction from separate events OpenTelemetry traces

The final row defines the benchmark question rather than expressing a generic observability preference: starting from the reported symptom, can an analyst select the first failed model, retriever, MCP, or tool invocation within the prescribed action budget and justify that selection without stitching together independent documents? Classify the symptom first. An error, timeout, retry, or bad dependency status makes the OpenTelemetry span canonical. For fully instrumented technical failures, parentage, status, exception events, and retry links make traces the overall winner for call-level isolation.

Once the first failed span is selected, attach structured logs as subordinate leaf evidence: a server error body, emitted tool arguments, a secure prompt reference, retrieval output, or validator result. A late top-level exception is a symptom of the failed workflow, not permission to relabel the final handler as the first fault. Keeping the span canonical preserves the distinction between the failure’s origin and its downstream propagation.

This also kills the free-text myth. A longer agent log does not reconstruct causality automatically. Without W3C trace context and stable call and attempt identifiers, even richly structured events remain a pile of correlated symptoms whose apparent order can change under concurrency. A retry line may be duplicated or interleaved; a vivid free-text exception can still point downstream rather than to the originating invocation.

Semantic diagnosis is the deliberate exception. Suppose the model, retriever, MCP server, and tool all complete technically, but an independent task oracle rejects the selected evidence, argument, or policy decision. The span tree proves sequence, not correctness. Here, the linked structured-log verdict wins because it records the oracle’s reason and identifies the bad semantic object; trace_id and span_id keep that verdict attached to the relevant execution.

For example, OpenTelemetry can show a retriever span completing before an MCP tool span returns a dependency error, making the tool the first failed invocation; a linked log can then supply its server body. Reverse the case: every span completes, but the validator rejects the retrieved document or tool argument. The execution trace remains correct, and the redacted validator event becomes canonical.

Treat this as a mechanism-backed decision rule, not empirical settlement from the supplied research. Quesma’s OTelBench item calls the benchmark independent, but the supplied description is a press release rather than evidence of peer review; the broader fetched record supports neither a traces win nor a logs win. The concrete next action is to score fully instrumented failures by symptom class, record analyst actions, and preserve the selected first-fault span together with its linked leaf evidence.

First-Fault Scorecard — Locate Failed Calls

What the Data Doesn't Tell You

The benchmark establishes an advantage only inside its observability envelope. Missing telemetry, biased retention, or changed schemas can turn a conditional result into a universal claim. According to Quesma’s OTelBench press release, issued in January of this year, the benchmark provides no trace-only versus log-only comparison, no agent-runtime failure corpus, and no per-signal ranking. For technical transport or dependency failures in the fully instrumented population, OpenTelemetry remains the winner; none of these caveats licenses switching the canonical signal to logs. The evidence gap limits what the result licenses us to conclude, not the symptom-first decision rule.

Consider a retriever–MCP–tool chain in which the retriever emits no span, a queued retry is manually parented to the model call, and host clocks disagree. A tool exception then looks both first and causal, although it is merely the next visible event. That trace does not identify the first fault; it contains a visibility gap followed by an ordering artifact. OpenTelemetry remains the canonical representation of the technical symptom, but the evidence may be insufficient to isolate the first fault. The benchmark’s isolation claim does not extend beyond its fully instrumented population. Cross-checking deployment topology and synchronized timestamps turns nesting into a testable hypothesis.

A longer free-text agent log cannot rescue that causal claim automatically. Without W3C trace context and call/attempt IDs, it is a pile of correlated symptoms. For a transport-clean semantic failure, the redacted structured event remains canonical and must link by trace_id and span_id. Still, generated rationale can rationalize a wrong choice, redaction can remove the decisive argument, and a task validator can encode an incorrect policy. An independent oracle should adjudicate the outcome, with its blind spots disclosed rather than hidden behind automation. The publication gate is a rerun with every disclosure below, not a pooled average that conceals stratum reversals.

Stress test How the apparent winner can be wrong Required disclosure
Span closure Trace-only analysis loses whenever any hop is uninstrumented. A missing retriever, MCP server, or tool span can make a later exception appear causal. Publish a span-coverage histogram and report missing-span cases explicitly.
Retention Head sampling can discard the incident entirely; tail policies improve failure yield but add decision latency. Report the retained-failure denominator separately from the number of displayed errors.
Version drift Changes in OpenTelemetry GenAI or MCP attribute names, span-status rules, and Collector defaults can make a newer fleet look better than an older one. Pin and publish every semantic-convention and Collector version.
Nesting Asynchronous queues, background retries, clock skew, and manually created parent links can distort the apparent order of model, retrieval, and tool calls. Cross-check the trace against deployment topology and synchronized timestamps.
Semantic oracle Agent-generated rationale can rationalize a wrong choice, redaction can remove the decisive argument, and a task validator can encode an incorrect policy. Adjudicate semantic failures with an independent oracle and disclose its blind spots.
Heterogeneity Model, tool, domain, temperature, load, and failure co-occurrence can change the relative result. Report stratified confidence intervals; if the ranking reverses in a material stratum, state that there is no universal winner instead of hiding the variance.
What the Data Doesn't Tell You — Locate Failed Calls

How to Choose Well

Choose by failure class, not by whichever record is more verbose. An error, timeout, retry, or non-2xx dependency status is a transport or causality failure: open the OpenTelemetry trace first and make its earliest failed attempt canonical. When every relevant model, retriever, MCP-server, and tool call completes but a task validator rejects the prompt, retrieval, arguments, or output, the fault is semantic rather than transport-level. Make the redacted structured log event canonical and retain the OpenTelemetry trace as sequence context. Classify the symptom before choosing the evidence source.

For example, suppose a retriever times out, its caller enters retry, and an MCP tool later returns a bad dependency status. Start with OpenTelemetry, then use parentage, operation links, and attempt IDs to determine whether the retriever attempt is causally upstream; its timestamp alone cannot settle the question. By contrast, if the model, retriever, MCP server, and tool spans all complete but the validator rejects the tool arguments, the redacted structured event carries the decisive semantic evidence. Link it to the relevant span through trace_id and span_id; neither an earlier span nor a longer log automatically wins.

A longer free-text agent log does not reconstruct causality automatically. Without W3C trace context plus call and attempt IDs, it is a pile of correlated symptoms. If any hop lacks a span or propagated trace context, label the root cause unknown and repair the instrumentation path. An absent record is not evidence that the hop succeeded, and an uncorrelated log event cannot be promoted to canonical evidence. This prevents a clean narrative from concealing a broken causal chain.

When diagnosis needs both topology and payload, preserve their division of labor. Put status, schema, token counts, and object references in spans. Keep redacted prompts and results in access-controlled structured log bodies. Join the representations through trace_id and span_id, but do not copy full payloads into span attributes. The result exposes the execution route and the decisive semantic artifact without turning sensitive prompts or results into broadly propagated telemetry.

RuleConditionCanonical option and required action
1Any model, retriever, MCP-server, or tool call errors, times out, returns a non-2xx status, or enters retry.Open the OpenTelemetry trace first and mark its earliest failed attempt canonical.
2Every relevant technical call completes, but the validator rejects the prompt, retrieval, arguments, or output.Mark one redacted structured log event canonical; retain the trace as sequence context and link both correlation identifiers.
3Any hop lacks a span or propagated trace context, including W3C context or call and attempt IDs.Label the root cause unknown, repair instrumentation, and reject both inferred success and uncorrelated free-text evidence.
4Multiple calls fail and their causal order is either provable or unprovable.With parentage, operation links, and attempt IDs, select the causally upstream failure; without proof, report the failure set rather than choosing by timestamp.
5Both topology and payload evidence are required.Keep status, schema, token counts, and object references in spans; keep redacted payloads in access-controlled structured bodies; correlate without duplicating full payloads.

What to do next

StepActionWhy it matters
1In the failed-call incident record, classify errors, timeouts, retries, and bad dependency statuses as OpenTelemetry-canonical. For a transport-clean but semantically wrong prompt, retrieval result, tool argument, or output, make the redacted structured agent-log event canonical and link both records by trace_id and span_id.A successful transport response does not prove that arguments, schema, state transitions, or business intent are correct.
2Represent each agent turn as an OpenTelemetry trace, with the agent root span parenting model-generation, retrieval, MCP client/server, and external-tool spans.A complete causal chain reveals where a failed call sits and prevents the investigation from starting at the final symptom.
3Pin the deployed OpenTelemetry GenAI and MCP semantic-convention versions in the instrumentation manifest, then verify parent-child propagation across every service and agent hop.Propagation—not span creation—was OTelBench’s most prominent failure mode; unpinned conventions can also change the fields used by isolation queries.
4For an OpenTelemetry-canonical failure, put Takeaway Detail Tracing in charge of walking the complete root-to-leaf topology and marking the earliest child that errors or times out—not the earliest log line or last successful retry.This locates the first causal edge requiring inspection rather than the loudest or most recent event.
5After identifying that edge—or immediately for a log-canonical symptom—open the redacted structured event linked by trace_id and span_id, then compare its payload and validator evidence.Logs explain prompt, retrieval, tool-argument, output, schema, and state-transition mismatches that trace status alone cannot explain.
6Record the benchmark caveat in the runbook: the leading model passed 29% of OTelBench instrumentation tasks but scored 80.9% on SWE-Bench; prohibit using that split to rank traces over logs.The benchmarks measured different work. Quesma reports that 39% of organizations cite observability complexity as their top obstacle and cites an average enterprise-outage cost of $1.4 million per hour, making coherent first-fault localization the priority.

Frequently Asked Questions

For a transport-clean call whose prompt, retrieval, tool arguments, or output is semantically wrong, which record should be canonical?

Make a separate redacted structured agent-log event canonical and link it to the successful invocation with trace_id and span_id.

How should the first failing edge be selected when retries and concurrent calls make wall-clock order misleading?

Follow parent-to-child causal edges and select the earliest child with ERROR or timeout status on each relevant route, not the earliest log line or final successful attempt.

What context must cross model, retrieval, MCP, and tool boundaries for a log event to identify one invocation?

Carry the 16-byte trace ID and 8-byte span ID across every boundary and copy both into the OpenTelemetry Log Data Model record.

Why does the leading model’s 29% OTelBench score not prove that traces outperform agent logs?

The same model scored 80.9% on SWE-Bench, but the benchmarks measured different work and did not compare traces with agent logs.

Why is accurate first-fault isolation an operational priority?

Quesma reports that 39% of organizations rank observability complexity as their top obstacle and cites an average enterprise-outage cost of $1.4 million per hour.

What must ATR-26 demonstrate for its stated thesis to pass?

It must resolve the first fault in no more than two analyst actions, exceed structured-log-only accuracy by at least 10 percentage points, and keep p95 overhead within the preregistered limit, with all three gates required for a pass.

Quick answers

What should tracing own when locating a failed call?Tracing should own first-fault localization by identifying the earliest causal edge needing inspection.
How is the first failing edge defined?It is the earliest child on a root-to-leaf causal route that errors or times out, not the earliest log line or final successful attempt.
Why is context propagation more important than merely creating spans?Emitting spans does not preserve execution continuity unless every hop accepts and forwards W3C Trace Context.
What canonical telemetry should represent a technical invocation failure?Use an OpenTelemetry ERROR span with an exception event, linked by the 16-byte trace ID, 8-byte span ID, and parent edge.
What should be recorded when transport succeeds but the result is semantically wrong?Emit a redacted structured agent-log event linked to the successful invocation by trace_id and span_id.

Also worth reading: From simple chains to interlocked workflows: a practical migration guide: From simple chains to interlocked · Agent tool failure recovery: 95% success with retry-first vs replan 2026: Agent tool failure recovery: 95%

Research Methodology & Editorial Standards

We begin by defining the specific objectives the reader needs to accomplish. Primary product documentation and authoritative secondary sources are assembled into a verified research corpus; drafting occurs only after this foundation is in place.

Every quantitative claim is subjected to dual-source verification. Any figure that cannot be independently corroborated is either qualified or omitted.

Published · Last reviewed · Owned by the Tryinterlock editorial desk (About, Contact, Privacy).

Related answers