The Chart is Not the Patient: Benchmarks Underrate the Evidence Problem in Clinical AI
David Vidmar ⋅ Nicole Hakim ⋅ Greg Lee ⋅ John Pfeifer ⋅ Logan Brigman ⋅ Christopher Haggerty ⋅ Brandon Fornwalt
Abstract
Agentic artificial intelligence (AI) holds great promise in chronic disease management, but will require robust evaluations to prove safety and effectiveness before being deployed in the clinic. Current evaluations rely on benchmarks largely focused on how well a system reasons $\textit{from}$ the record. We argue that agents must also reason $\textit{about}$ the record: judiciously deciding what to trust, ignore, or escalate in electronic health records (EHRs) that are uniquely and inherently messy. We call this second layer “clerical uncertainty”, to distinguish it from clinical uncertainty about a patient's state, and show that three generations of benchmarks have progressed primarily along the clinical axis while leaving clerical uncertainty under-measured. Worse yet, we argue that handling clerical uncertainty is exactly where large language models are known to be weak. Finally, we define a defect-response grid that can serve as a target for the more comprehensive evaluations needed for agentic AI to safely reach the frontlines of medicine.
Chat is not available.
Successful Page Load