Evaluation format, not model capability, drives measured triage failure in the assessment of consumer health AI
David Fraile Navarro ⋅ Jialei Sheng ⋅ ⋅
Abstract
A recent Nature Medicine study reported that ChatGPT Health under-triages 51.6\% of emergencies and concluded that consumer-facing AI triage poses safety risks. Its protocol, however, was an exam-style scaffold (forced A/B/C/D output, knowledge suppression, no clarifying questions) that differs fundamentally from how consumers use health chatbots. We ask whether the headline reported triage error rate is a property of the models or of the measurement. In a first, mechanistic study, five frontier LLMs on a 17-scenario bank with similar questions, scored 6.4 points higher under naturalistic patient-style messages than under the constrained scaffold ($p=0.015$), and a sensitivity analysis of the asthma vignette showed three models scoring 0--24\% with forced choice but 100\% with free text. In a second, faithful replication of the original study we ran the authors' own 60 released reference vignettes through six frontier models under four matched formats, with clinician validation of the naturalistic rewrites and a blinded clinician audit of the LLM adjudicators. Here the direction reversed: information-preserving free-text rewrites scored slightly below the exact structured prompt (78.7\% vs 81.8\%, $p=0.020$), removing only the answer scaffold from the original wording changed aggregate accuracy little (80.6\% vs 81.4\%, $p=0.81$), and naturalistic input with a forced categorical answer beat both the exact prompt (84.4\% vs 81.8\%, $p=0.023$) and free text (84.4\% vs 78.7\%, $p=0.0001$). On the four vignettes that define the original study's emergency rate, under-triage was 17\% with the exact scaffold, 50\% when the same cases were answered in free text, and 21\% when the same naturalistic message was answered with a forced letter; every free-text ``under-triage'' was a same-day recommendation scored as C rather than D, and 60\% of free-text answers made escalation conditional on information the patient was asked to check or report (which is how turn-based chatbots operate), the interactive behavior a single-turn benchmark cannot score. The headline rate therefore did not reproduce as a stable cross-model finding; it is largely a property of the output format and of mapping prose back onto a four-point scale. Benchmark scaffolds for consumer health AI are behaviorally active instruments, and safety claims should report sensitivity to input wording, output format and adjudication.
Chat is not available.
Successful Page Load