Sauron: An Evidence-First Agentic Verifier
Abstract
Long-horizon agent tasks are increasingly evaluated with rubrics rather than deterministic tests, and rubric verdicts are beginning to serve as training rewards. Many existing verifiers judge from the agent's final response or serialized trajectory, treating the presented account as reliable even though the evaluated agent controls much of it. We present Sauron, an evidence-first agentic verifier that inverts this relationship. An investigator interrogates the final environment state, artifacts, and execution trace through bounded queries. A claim can support a verdict only after its source is resolved and its support verified. Exact claims are checked mechanically; semantic claims undergo an isolated faithfulness check that never sees the investigator's reasoning. State is authoritative for outcome claims; the trace provides investigative leads and process evidence but cannot establish outcomes on its own. Every verdict includes a receipt identifying the evidence on which it rests. On 2,586 human-labeled criteria from 100 internal, long-horizon personal-assistant episodes, Sauron achieves the highest F1 of five verifiers on all eight backbones tested. It costs less than Gandalf, the strongest baseline, on every backbone while processing roughly half as many input tokens, and its lead widens on process criteria. On a paired fault-injection suite spanning 13 families, Sauron achieves a detection-minus-false-alarm margin of 0.950; trajectory-blind verifiers remain below 0.1, while full-trace-as-text variants reach at most 0.714. Ablations show that trace access and evidence admission drive quality, while bounded navigation and clustering drive efficiency.