Invocation-Level Reliability of Tool-Using Agents
Abstract
An early tool-call error can corrupt every downstream input even when an agent subsequently applies the right local rule. We isolate this effect with matched evaluations under a correct baseline history and the model's own free-running history. At each task depth, the pooled invocation rates in these two conditions define a fit-free relative propagation loss: one minus the free-running rate divided by the baseline rate. Across five open-weight models on controlled routing tasks, the two 7-8B models lose about 68% of baseline capability by depth 6; stronger models mostly remain at ceiling. Our main result is a construct-validity condition for trajectory scoring. If the canonical next target after a divergence has conditional guessing probability at most epsilon given the agent-observable history, then any fixed-gold scorer credits the agent with probability at most epsilon, regardless of its local competence. Gold-agreement severity is therefore driven to a scorer-imposed boundary, and observed reconvergence cannot measure behavioral recovery.Fixed reference alone is not sufficient: the result applies when the canonical target becomes unreachable or unpredictable, and excludes reconstructible targets and evaluators that accept alternative states or milestone paths. Our hidden-constant tasks instantiate the condition: 0 of 869 corrupted-context calls match gold, and 0 of 580 opportunities reconverge. Conditional-on-state replay instead asks whether a call correctly continues from the state the agent actually holds; it yields interior severity estimates of 0.149 and 0.316 and requires no new model queries. Alternative absolute-gap and error-amplification summaries give the same qualitative depth trend.