Position: Health Chatbots Must Be Evaluated Across the Disclosure-to-Action Pathway
Abstract
Evaluations of patient-facing health chatbots usually begin with a complete medical question, which omits the earlier interaction in which a user may hesitate to describe a sensitive symptom. In testicular cancer, embarrassment has been linked to delayed help-seeking, and diagnostic delay has been associated with more advanced disease. We argue that evaluation should instead cover the full disclosure-to-action pathway, from a user's initial concern to the recommendation for care, rather than maximizing how much a user discloses. We define decision-relevant, minimum-necessary disclosure as the target, whether a chatbot obtains the information needed for a safe recommendation without asking unnecessary sensitive questions, and specify which of these claims scripted tests can support and which require studies with human participants. Safety failures such as under-triage or coercive questioning should be reported as disqualifying rather than averaged into a composite score. To test whether current practice already covers this pathway, we audited five testicular-cancer chatbot evaluations and found that all five used complete single-turn prompts, with none testing how a chatbot responds when clinically relevant information has not yet been disclosed. In a probe of three openly available assistants across twelve cases in both forms, replies to an incomplete opening reassured the user before establishing the relevant fact in 20 of 36 cases, against 6 of 36 when the symptom was stated in full. Evaluation that begins only after a concern has been fully stated cannot show whether a chatbot safely supports the harder, earlier part of the conversation.