Before Blaming a Call: Three Validity Gates for Agent Failure Attribution
MARCEL OSMOND ⋅ Kenzhebaev Bektur ⋅ Jinming Zhang
Abstract
Agent-failure attribution benchmarks often score whether a localizer names an injected call and whether replacing it repairs the outcome. Such scores can be precise when the target itself is not empirically identified. We propose three prerequisite gates: the injection changes an independently scored outcome (G1); the declared trace representation supports origin ranking beyond ordinary variance (G2); and repair evidence separates intervention from stochastic reruns and tests joint causes (G3). We audit one fixed six-stage topology under two repository-created controlled tasks: VANTOR, which propagates a fictional numeric survey value, and Meridian, which propagates a fictional categorical eligibility mapping. In VANTOR, 46/156 single-fault traces remained correct; among 36 open-weight incidents, exact ranking changed from 15/36 with lexical traces to 31/36 with task-aligned fields, while a fixed-position prior scored 36/36. In preregistered Meridian pairs, Qwen showed 18 observed conversions versus 2 fault-helped outcomes ($\Delta=.80$, $p=.00040$), whereas Mistral showed 6 versus 6 ($\Delta=0$, $p=1$); replicated G1 support is therefore mixed and the formal G2 comparison remains gate-closed. VANTOR state replacement often restored outcomes, but no matched no-repair control was collected, so G3 is not established. In separate double-fault trials, no singleton repair was observed to suffice in 11/20 scored cases. We claim neither a new localizer nor unique causal attribution: accuracy should be conditioned on validated incidents, representations, and interventions.
Chat is not available.
Successful Page Load