Brittle Trust, Recoverable Control: Robust Evidence Arbitration in Tool-Using LLMs
Abstract
Tool-using agents must arbitrate conflicting evidence as tool reliability changes. We introduce a controlled robustness map that varies matched tool histories, dynamic reliability shifts, and three held-out natural-language tool surfaces while holding current evidence fixed. Qwen2.5-7B-Instruct uses observed reliability history on the base surface (73.96\%, 95\% CI 71.25--76.57), yet falls to chance on arithmetic, calendar, and database surfaces and exhibits a sharp degradation--recovery asymmetry after unannounced reliability shifts. Reliability remains linearly decodable, and a residual direction learned only on the base distribution transfers zero-shot across all three surfaces: signed by a fixed empirical-history rule, it raises 7B from 50.0\% to 100.0\% across 480 held-out decisions. For 3B, the same intervention transfers only partially, improving pooled surface performance from 59.4\% to 67.1\%. Together, these results expose a representation--behavior gap: reliability can remain internally accessible even when native evidence arbitration is brittle, while the same representation can still support recoverable control.