TIB: Sample-wise Tempered Information Bottleneck for Multimodal Attribution beyond Alignment Assumption
Abstract
Multimodal (e.g. vision–language, in this work) attribution aims to interpret the models with- out access to ground-truth explanatory supervi- sion. Existing multimodal attribution methods typically rely on paired modalities as proxy su- pervision, implicitly assuming that image–text pairs are semantically aligned. Information bot- tleneck (IB)–based attribution methods follow this general paradigm and instantiate it through a global sufficiency–compression objective. In real- istic multimodal datasets, however, this assump- tion is frequently violated due to partial semantic misalignment and unreliable cross-modal corre- spondences. Through theoretical analysis, we show that, under such misalignment, information- bottleneck-based attribution objectives exhibit a structural limitation: by assuming equal reliabil- ity across all proxy supervision signals, they in- duce posterior over-contraction and forced expla- nations. To address this issue, we propose TIB, a sample-wise Tempered Information Bottleneck framework that adaptively modulates attribution strength according to cross-modal reliability, with- out assuming semantic alignment. Extensive ex- periments on controlled benchmarks and large- scale datasets demonstrate the effectiveness of the proposed approach, while also validating the identified failure mode