RankAlign: Unsupervised Vision-Language Representation Alignment via Rank Transformation
Abstract
The Platonic Representation Hypothesis suggests that vision and language models converge toward a shared latent structure, yet the unsupervised alignment of unpaired modalities remains a fundamental challenge. Current methods rely on direct comparisons of intra-modality similarity kernels, which are often incommensurate due to disparate scales and representation densities. This mismatch yields ill-conditioned optimization landscapes, hindering stable alignment. To address this, we propose RankAlign, a framework that independently transforms each within-modality kernel into a rank-based representation of relative neighbor relationships. By shifting from absolute similarity values to ordinal relational geometry, RankAlign ensures structural consistency while remaining robust to modality-specific distortions. We show that rank-based transformation reshapes similarity distributions, mitigating representation collapse and providing a more discriminative alignment signal. Experiments demonstrate that RankAlign is noise-resilient and overcome existing kernel-based alignment method across diverse datasets, including a medical imaging dataset where subtle class differences typically limit existing approaches alignment approaches.