CoMorBench: Language Models Fail to Track Simultaneous Hypotheses under Interaction
Abstract
Clinical language-model benchmarks typically assume a single correct diagnosis with the evidence provided up front. However, real patients present with several diagnoses. Symptoms must be gathered over a consultation. We study diagnoses where these two demands meet, multi turn and multi label at once. We introduce CoMorBench, a benchmark in which a doctor model interviews a patient whose ground truth is a set of co-occurring diagnoses. Its cases come from a generation algorithm that draws realistic multimorbid vignettes from the ePOCT+ pediatric decision tree with a comorbidity model calibrated on real consultations. Across six models, the metrics that are strong on single-diagnosis cases collapse once diagnoses co-occur under interaction. With our experiments, we showed that models, in interactive settings, act as single hypothesis trackers that recover each disease in isolation but lose it when another shares the context. We support this finding with empirical and theoretical analysis.