Beyond the evidence boundary: Behavioral Mirage Under Progressive Degradation in Medical VLMs
Abstract
VLMs for medical image interpretation are typically evaluated on clean inputs, despite the progressive evidence loss that can occur in high-stakes clinical settings. Existing robustness studies show that performance can degrade while confidence remains high, but do not establish when a previously reliable model--case pair ceases to behave reliably. We introduce a baseline-anchored framework that traces verified-correct cases through progressive image and language degradation to identify an empirical evidence boundary, and define mirage as continued commitment past this boundary, regardless of correctness. Across three VLMs and three medical datasets, only MedGemma provides a sufficiently large baseline-correct pool for trajectory analysis. Its post-boundary behavior is heterogeneous, including direct mirage, correct recovery, mirage dissolution, consistent abstention, and oscillation, with mirage rates of 16.0--100.0% across dataset--modality pairs. Language-only answerability strongly predicts baseline correctness but provides limited discrimination among post-boundary regimes (AUC 0.618). Despite this behavioral heterogeneity, residual-stream probes decode eventual commit-versus-abstain behavior with AUROC of 0.788--1.000. These results show that progressive evidence loss can produce persistent commitment poorly reflected by stated uncertainty, motivating boundary-aware evaluation and improved abstention mechanisms for reliable medical foundation models operating under partial observation.