DropBench-MM: How Do Multimodal Conversational Models Degrade When Modalities Disappear?
Md Rashad Al Hasan Rony
Abstract
Deployed systems lose modalities constantly, whereas multimodal leaderboards report one number measured with every stream present. A model selected on that basis therefore comes with no evidence about the conditions it will meet. We introduce DropBench-MM, an inference-only benchmark that scores frozen checkpoints under seven deployment-realistic corruptions, from whole-modality loss to simulated ASR error, at up to four severities each. Three commitments fix what the instrument measures. First, all randomness is keyed to the corruption specification and the segment, making any evaluation independently reproducible. Second, every model scores the same materialized batch, and none sees different input. Third, each corruption maps onto a common damage axis and is read as \emph{retention} against the model's own clean score and averaged over types into a single DropScore. Applying it to 11 models across 8 fusion families and three corpora yields two kinds of result. About the models, clean accuracy predicts robustness on none of the three corpora. On CMU-MOSEI it anti-predicts robustness at $\rho = -0.727$, and the reason is identifiable rather than arbitrary rank fluctuation, since the clean leaders are the two architectures with no cross-modal interaction during encoding. The robustness ordering itself does not transfer across corpora. About the instrument, three of the seven types prove inert on word-aligned features, a property of the feature release rather than of the corruptions. Two of four validity checks fail, one of them a ranking that depends on the fill policy. We release the harness and per-sample predictions for every evaluation, together with the reporting protocol these results imply.
Chat is not available.
Successful Page Load