Verify Claims, Not Scores: Diagnosing Modular Agents
Abstract
Developers typically judge a change to one component of an agent, such as its controller, a learned model, or its verifier, by whether an aggregate task score improves. Yet a flat or rising score cannot say whether improvement was attainable, which component lost value, or what the agent's own checks actually certify. We argue that the unit of verification should be the claim, not the score, and introduce a claim-specific verification audit for modular agents that plan, act, check, and refine. Every conclusion is recorded with its evidence, a verdict (supported, unsupported, unresolved, or not evaluated), and the boundary within which it holds. The audit targets three ways in which an aggregate score misleads and pairs each with a diagnostic: the evaluation may be unable to express the improvement, tested with oracles over nested action sets; a downstream component may mask an upstream one, tested with one-at-a-time oracle replacements whose null results are treated as unresolved; and a verifier may not identify the quantity it is read as certifying, tested by examining what its score can and cannot measure. Applied to a constrained portfolio-allocation agent in a synthetic market with known hidden regimes, the audit shows that the value of perfect regime information depends on the action set used to measure it, that the scenario generator discards most of the regime signal while better local fidelity does not improve decisions, and that the runtime verifier can be bypassed with no visible change in outcomes. The contribution is the audit protocol; the empirical findings are specific to the agent and environment studied.