Calibration Collapse in Multi-Turn Diagnostic Dialogue with Large Language Models
Aarav Agarwal ⋅ Abhiram Kuchi ⋅ Kishore Nuthalapati ⋅ Keyi Xue ⋅ Kiran Nijjer
Abstract
Large language models (LLMs) deployed as diagnostic dialogue agents must be well-calibrated, since clinician trust hinges upon stated model confidence. We evaluate 214 AgentClinic-MedQA cases across four backbones (GPT-4.1, GPT-OSS-120B, and DeepSeek-v4-flash-0731 in reasoning and non-reasoning modes) under eight conditions that hold task content fixed while varying turn count, evidence order, and interactivity. Calibration collapse is most prominent in GPT-OSS-120B, in which accuracy collapses from 88.8\% (static) to 38.3\% (sequential) even though the model's confidence drops only from 88.1\% to 72.8\%. Statically re-presenting the same sequentially-gathered evidence yields similarly low accuracy (33\%--43\% vs.\ 79--89\%), suggesting the collapse is not an artifact of the prompt's format itself. Case-level analysis shows oracle and live-sequential runs fail on the same cases (Cohen's $\kappa = 0.69$--$0.90$, all McNemar tests non-significant, $p = 0.34$--$1.00$). Furthermore, incremental presentation---identical turn count and case content, but no self-directed questioning---succeeds on 95--105 of the cases where sequential dialogue fails (vs.\ only 3--5 the reverse direction), with low case-level agreement ($\kappa = 0.126$--$0.191$) and all four McNemar tests decisively significant ($p < 10^{-20}$). Confidence nonetheless retains meaningful rank-correlation with correctness throughout (point-biserial $r$ up to 0.62, $p < .001$) and providing models with pertinent feedback history substantially improves calibration across all backbones (ie. GPT-4.1 ECE falls from 0.301 to 0.176). These results indicate that calibration collapse reflects models' failure to discount confidence for incomplete self-gathered evidence, not an inherent inability to recalibrate---self-directed evidence acquisition simply never supplies the corrective signal that would allow it.
Chat is not available.
Successful Page Load