Observability Debt in Agent Evaluation: Toward a Telemetry Layer for AgenticOS
Abstract
A resource-aware system layer for agents needs one joined record of what ran, on what, and at what cost. Today that record is broken across three places that never meet: a leaderboard, a trace corpus, and a telemetry log. None of the sources we examine keeps in one place the fields such a layer would need to attribute cost: task, model, scaffold and policy configuration, outcome, token and cache usage, retries, latency, and billing. We study what this missing record costs, drawing on three public datasets. With the Holistic Agent Leaderboard (HAL), raw dollar- cost orderings do carry across some benchmark pairs (Spearman ρ=0.74–0.93) but are largely explained by a persistent displayed-label cost tendency (pooled R2=82.0%); they weaken sharply once we adjust for success (mean ρ=0.59 against 0.84). With 977 task- and model-matched trajectory pairs from Open- SWE-Traces, scaffold choice shifts measurable execution burden at fixed task and model (OpenHands versus SWE-agent turn ratio 0.85, interval 0.82–0.87), and the common belief that failures are far more expensive is not borne out under exact task matching (our design has 80% power for premium fractions at or above 0.65; we observe 0.44–0.49). Across two independent telemetry corpora (8,058 and 1,017 sessions) we document the resource profile a scheduler would need but that no benchmark links to an outcome: a 93.7% mean cache-read share, tool-error cascades 4.29× more likely after an error, and the top 1% of sessions holding 48.3% of input-token volume. The three studies complement rather than combine. Studies 1 and 2 show that raw-cost orderings and naive failure premiums mix up which workloads were attempted with how agents behaved, and Study 3 shows the execution variables hidden under those totals are themselves uneven, heavy- tailed, and state-dependent. From these observations we derive a joint telemetry schema spanning identity, execution, resources, control, and outcome, and argue it belongs at the operating-system layer, because the layer that allocates resources is the one that must observe what those allocations depend on.