Questions Before Answers: Rethinking Safety Evaluation for Patient-Facing Health AI
Abstract
Patient-facing health large language models (LLMs) are commonly evaluated by giving them a complete clinical vignette and scoring the diagnosis, triage category, or recommendation they produce. We argue that this evaluates the answer while assuming away a capability that safe patient-facing triage requires: asking for information the patient has not provided. A patient may describe a headache without mentioning sudden onset, neurological symptoms, or anticoagulant use, and a safe system must recognise what is missing before deciding what the patient should do. Yet a model that asks no useful questions can score identically to one that elicits the clinically decisive information first. Evidence identified during a prospectively registered systematic review of patient-facing LLM safety suggests the distinction is consequential. Three models differing in classification accuracy by fifteen percentage points omitted the same safety-critical question; a model became 16% as likely to sustain dialogue by the seventh turn of an escalating risk disclosure; patients supplied reports 8% less suitable for urgency assessment when they believed the recipient was an AI; and a system with zero single-turn errors produced 40-50% error rates once the conversation continued. We propose three additions to current benchmarks: required-question coverage, missing-information pairs, and turn-budget safety curves. As health systems become conversational and agentic, deciding what to ask next becomes part of the action policy.