Trajectory-Derived Confidence for Reliable, Resource-Aware Clinical Text-to-SQL Agents
Mincheol Song ⋅ Joshua Ward ⋅ Jake Jung ⋅ Guang Cheng
Abstract
LLM agents for clinical text-to-SQL applications reason autonomously over multiple steps but cannot assess whether their own reasoning or outputs can be trusted. In high leverage applications such as healthcare, this presents a critical risk where system mistakes can be costly. These reliability failures are also resource failures: an incorrect reasoning trajectory spends computation budget on outputs that must be discarded. We introduce Sentinel, a trajectory-derived, classifier-based confidence layer that analyzes an agent's reasoning, code and database outputs to decide at three points whether to stop: refusing unanswerable questions before the agent runs, halting doomed trajectories mid-run, and withholding untrustworthy answers at delivery. Here, utilizing Chow’s rule, we optimize decisions under EHRSQL’s Reliability Score, which penalizes incorrect answers given a utility weighting, and find on the benchmark EHRSQL that Sentinel improves this scoring by 200\% when mistakes have a low utility weighting and turns a strongly negative score positive ($-3.02$ to $+0.03$) at higher stakes. We further find that delivered-answer accuracy rises from 54\% to 69--84\% as the cost of mistakes increases, creating a lower risk to patients. These three stopping modes also substantially improve the computational cost of agents where we find Sentinel eliminates up to 78\% of all agent steps, a measured $5.0$ seconds of latency per question, with zero reliability cost.
Chat is not available.
Successful Page Load