Lengthy Road to Uncertainty: Estimating Confidence of Black-Box Agentic Text-to-SQL
Abstract
Verifying whether an agent has completed a task correctly is a prerequisite for reliable agent development. We study this for text-to-SQL agents, where an incorrect query typically returns a plausible, yet conceptually wrong table rather than raising an error. Assuming the agent cannot be rerun and its internals are unavailable, we compare four verifiers of a completed trajectory across three agents and three benchmarks: an LLM judge, the agent's self-report, uncertainty estimators from a small surrogate model teacher-forced over the saved trajectory, and trajectory length. Surrogate-based estimators reach 0.752 AUROC but are consistently outperformed by the judge and do not improve with surrogate scale, while a weaker judge performs near chance. Trajectory length aligns with most surrogate estimators, but it proxies for task difficulty rather than correctness and should be reported as a baseline in agentic verification. Combining the judge with trajectory length is strongest (0.864 AUROC), and verifier choice is governed by cost and access rather than accuracy alone.