When Can You Trust an Agent’s Own Verifier? Stress-Testing Counterfactual Rationale Verification Under Search Pressure and Oracle Access
Abstract
Automated agent design increasingly trusts a single score to decide whether a change is an improvement — but a score optimized aggressively enough ceases to measure what it was meant to. When an agent also explains its design, that explanation is a second signal we can check without human labels: negate each stated claim, re-run, and observe whether performance drops. Such verifiers are appealing, so we ask whether they stay reliable once an agent optimizes against them. Our contribution is not a new verifier but a way to stress-test one: a pressure ratio ρ measuring how much the verifier steers the agent's search, and four falsifiable probes, one per failure mode — giving no signal, being too weak to matter, being gamed by stating fewer claims, and being queried before the agent commits. We run the test on one concrete verifier, the Rationale Consistency Score (RCS), inside an LLM agent that designs machine-learning pipelines — tool-using in the probe arm — across six tasks and three seeds. We report three findings. First, as deployed the verifier exerts negligible influence on the search objective (ρ ≪ 1), so its robustness under optimization had never been meaningfully tested, even though it still shapes the reported artifact through a final-selection tie-break that ρ does not register. Second, even after we raise the pressure more than two orders of magnitude above its highest deployed level, the search still does not discover the claim-minimization exploit — yet a simple instruction shows the agent could readily do so. The shortfall is thus in the search loop rather than the agent: feedback never exposes claim volume, so the reward's standing incentive to delete inert claims never becomes legible to search. Third, the failure that does appear is query access: when the agent can probe the verifier before committing, it keeps only the claims that pass, so its rationale appears verified even as its held-out faithfulness drops — a decline invisible to the in-loop metrics (non-overlapping CIs on one dataset, overlapping on two more). The practical implication is a reporting standard: audit any agent's verifier with ρ and these four probes before trusting it.