Who Verifies the Tools? How Warped Formatters Mislead Agents and Their Verifiers
Abstract
Agents and verifiers only ever see the text a tool prints, and that text can misrepresent what the tool actually did, even when the tool's computation is correct. We saw this happen in a production RAG research system serving a historical archive in New York, where tool output formatters misled an agent into giving wrong or incomplete information to users. Each tool did its job correctly. But the formatted output fed to the agent left out key details: that the search had covered only the first 1 million characters, that the result list had been cut at 20 items, that a number was an estimate rather than a recorded fact, and that an identifier came from a different numbering system. We studied whether these formatters misled agents and verifiers and how much the repaired wording helps. We fed tool results to agents using two formatter versions: one with the key omissions and one with all missing details clearly stated. That difference alone changed what the agents claimed and what the verifiers accepted. Unsupported agent claims dropped substantially for three of the four failure types. In one test, verifiers went from accepting incorrect answers 90 out of 90 times to rejecting them 90 out of 90 times. A control showed the verifiers could catch the errors once they were given the missing facts and asked to judge correctness. We include the audit prompt we used with Codex and Claude Code to find warped formatters like these. When pointed at an old version of our own system, it rediscovered the defects that were still live and surfaced twelve more we had not catalogued. When pointed at OpenClaw and Hermes, two of the best-known open-source agentic systems, it found warped formatters in both. We opened PRs to fix both, and the OpenClaw fix has already been merged. Practitioners can run the same prompt on their own systems today.