Who Verifies the Agents? Toward Reliable Agent Development
Abstract
Recent advances in autonomous agents have demonstrated potential for complex reasoning and open-ended tasks. However, current agent development is often fragmented and dependent on ad-hoc methods, leading to reliability challenges where performance plateaus or regresses during iteration. The fundamental bottleneck preventing the transition to scalable, reliable agent systems is verification: the inability to systematically determine whether a change to an agent constitutes genuine improvements. This workshop seeks to formalize verification as a core discipline in agent development. We convene researchers and practitioners to explore three critical pillars: developing robust verifiers that resist reward hacking, leveraging environment-grounded simulation as the ground truth for evaluation, and integrating heterogeneous signals like latency, cost, and calibration into the verification loop. By centering verification in the development lifecycle, this workshop aims to establish the foundation necessary to move toward truly autonomous and reliable agentic systems.