MedRouteBench: Controlled Evaluation of Evidence-Conditioned Revision and Stopping in Medical LLM Agents
David Man ⋅ Chuanhai Xu ⋅ Quang Tran ⋅ Joshua Liu ⋅ Ryan Bui ⋅ Benjamin Liu ⋅ Kevin Zhu
Abstract
Final-answer accuracy does not reveal whether a medical language model corrected an unsupported answer, changed a correct one, or stopped before useful evidence became available. We introduce MedRouteBench, a controlled evaluation of answer revision and stopping under staged biomedical and clinical evidence. We adapt 482 PubMedQA cases into a two-stage evidence-revision task and replay 107 MedCTA tool trajectories under stopping-allowed and forced-continuation conditions. We evaluate ten models with three runs per condition. With the complete PubMedQA abstract and revision prompt, final accuracy exceeded preliminary accuracy by 10.9--59.8 percentage points across models. However, successful revision among initially wrong answers ranged from 51.2--82.4\%, while overreaction among initially correct answers ranged from 5.7--51.4\%. Presenting the results context first increased preliminary accuracy by 9.3--59.7 points, whereas final-accuracy changes were smaller and mixed ($-4.7$ to $+9.4$ points). In MedCTA, every model declared a final answer before the recorded endpoint in at least half of the cases under both replay conditions. Thresholded final-answer accuracy was 6.2--29.6 points higher under forced continuation, which supplied the remaining reference observations after an attempted early exit. These results show that endpoint accuracy can obscure substantial differences in selective updating and stopping. MedRouteBench therefore reports correction, overreaction, and stopping as separate outcomes under fixed benchmark protocols.
Chat is not available.
Successful Page Load