On the Geometry of Multimodal Saturation: Riemannian VICReg
Nessim Ben Abbes ⋅ Duc-Han LE ⋅ Sabri Mtibaa ⋅ Van-Tam Nguyen
Abstract
In self-supervised learning, a third modality should improve, or at least preserve, performance. Across nine image-text-tabular datasets, we show that it instead harms performance: the trimodal model underperforms its own best bimodal subset in $55.6\%$ of paired runs under VICReg. The same failure occurs in $51.1\%$ of paired runs under SimSiam. We call this failure multimodal saturation. We propose that the failure lies in the alignment geometry. Riemannian VICReg (R-VICReg) generalizes classical VICReg: it aligns views by squared geodesic distance on learnable negative-curvature product factors and recovers VICReg exactly as curvature vanishes. Over the same $45$ paired runs, R-VICReg raises the probability that the third modality helps from $44.4\%$ to $64.4\%$, with gains concentrated where VICReg saturates.
Chat is not available.
Successful Page Load