No Free Alignment: Observability-Aware Alignment for Multimodal Heterogeneous Learning
Abstract
Multimodal heterogeneous learning aims to train reliable multimodal models when data sources differ in modality availability, semantic coverage, domain distribution, and missingness patterns. Although cross-modal alignment is widely used to mitigate such heterogeneity, we argue that alignment is not free: enforcing relations that are weakly supported by the observed data can amplify noise, over-transfer missing-modality information, and induce negative transfer. This paper studies when a modality--modality--semantic relation is sufficiently observable to be safely aligned and shared across heterogeneous sources. We propose ObsAlign, an observability-aware framework that summarizes visible modality--semantic evidence through compact source-level sketches and estimates relation-level support from bridge evidence, side support, cross-modal consistency, and prototype uncertainty. The resulting observability score gates co-observed alignment, calibrates weak missing-modality transfer, and guides evidence-aware source consolidation so that reliable relations are shared while unsupported variation remains local. Across image--text, audio--visual--text, action-recognition, sensor, and healthcare benchmarks under source-partitioned multimodal heterogeneity, ObsAlign achieves the best overall performance in both encoder-based and VLM-compatible regimes, while sharply reducing negative transfer on low-observability relations. These results suggest that relation-level observability provides a practical principle for deciding when cross-modal alignment should be strengthened, weakened, or avoided.