Auditing Trust Signals in Black-Box Medical VLMs: What to Perturb and What to Ignore
Baishali Chaudhury ⋅ Mohamed Zaid ⋅ Jae Oh Woo
Abstract
Deployed medical vision--language models are often accessed through APIs that expose no logits or internal uncertainty signals, leaving a site to judge reliability from model outputs alone. We study which black-box elicitation changes generate behavior that is useful for detecting prediction errors. We cross a direct-versus-deliberative reasoning-mode contrast with three response/output specifications in a $2\times3$ design and evaluate all two-query subsets across seven configurations spanning four base models on chest radiography and dermoscopy. Across both datasets, reasoning-mode contrasts are consistently more error-discriminative than the tested output-specification contrasts, while several chain-of-thought format comparisons produce no label variation at all. A two-query direct-versus-deliberative score retains $77\%$ and $88\%$ of the full six-query battery's above-chance discrimination on CheXpert and ISIC, respectively, and outperforms six-query stochastic resampling in the evaluated configurations. In selective-review simulations, the fused behavioral score removes $34\%$ (Chexpert) and $57\%$ (ISIC) of errors at $30\%$ deferral. By contrast, self-reported confidence can reverse orientation across models and tasks; its agreement with the behavioral score is strongly associated with its error-discrimination performance ($\rho = .79$ and $.89$ across evaluated configurations), suggesting an exploratory label-free diagnostic. These results support careful selection of perturbation type when auditing black-box medical VLMs and caution against treating stated confidence as a reliable routing signal without task-specific validation.
Chat is not available.
Successful Page Load