Sycophancy and Stubbornness in Medical Language Model Post-Training
Abstract
Medical language models should change diagnoses when clinical evidence changes, but not endorse unsafe patient beliefs under pressure. We study sycophancy, agreeing with patients' unsafe beliefs under pressure, and diagnostic stubbornness, giving the same diagnosis despite a change in clinical evidence. We evaluate ten open models, including Qwen3, HealthGPT-Pro, and HuatuoGPT-3, using MedEinst counterfactual cases and MedPRESS pressure dialogues. We find that models with low sycophancy can still exhibit diagnostic stubbornness. To test whether diagnostic supervision changes these behaviors, we train Qwen3-4B and HealthGPT-Pro-4B on identical filtered teacher demonstrations using supervised fine-tuning (SFT) and token-level distillation. Both methods improve paired diagnostic accuracy, with SFT achieving higher scores at matched steps. On test pairs whose original cases are correctly diagnosed before and after training, HealthGPT-Pro-4B reduces diagnosis repetition by 10.9--14.2 percentage points, whereas Qwen3-4B shows no statistically resolved reduction. Yet, HealthGPT-Pro-4B also becomes more likely to agree with unsafe patient beliefs. Overall, our results show that improved accuracy and reduced diagnostic stubbornness need not preserve resistance to patient pressure.