Recovery Is a Decision - But Whose? Live Agents Overturn Replay Evaluations of Tool-Failure Triage
Abstract
Tool-augmented LLM agents handle tool failures badly: they ignore error signals, retry the same failing call, and burn their step budgets. A growing line of work responds with recovery machinery, from fine-tuned recovery policies to hand-designed orchestration rules to prompted LLM supervisors. We study a far lighter alternative: a logistic-regression triage controller that watches a frozen agent for tool-call failures and chooses retry, switch, or abort. On 16,163 real failure events from PALADIN's ToolBench corpus, this simple learner agrees with annotated recovery decisions far more often than PALADIN's own retrieval mechanism or rule-based orchestration (63.6% vs. ~43%), and larger models add nothing. In scripted replay of SciAgentGym tool chains it also attains the lowest recovery-regret point estimate of the policies we test. But when we rerun the same evaluation with live agents in place of scripted replay (two open-weight backbones, ~1,900 episodes, the same failure injector), the result reverses under our injected-failure regime: every external triage policy, ours included, underperforms the agents' native error handling. Replacing abort with deferral to the agent recovers most of the deficit, and a sweep over injector parameters leaves the replay conclusions unchanged. Together these locate the flaw in the executor: a scripted replay never recovers on its own, so an abort looks cheap there, while in a live run it forfeits episodes the agent would have finished. Recovery-policy value is a joint property of the agent, the failure distribution, and the executor, and it must be measured against executors that can recover on their own. We release 6,007 live-harvested failure events and a fully cached, deterministically replayable evaluation.