Spherical Interpolation for Backward-Compatible Multimodal Representations
Simone Ricci ⋅ Niccolò Biondi ⋅ Federico Pernici
Abstract
Contrastive vision-language models map visual and textual representations in a shared normalized embedding space, making cosine similarity the natural metric for cross-modal retrieval. A practical challenge arises during model upgrades: independently trained models generally produce incompatible representation spaces, so replacing a deployed model typically requires recomputing embeddings for the entire gallery, which is prohibitively expensive at scale. Orthogonal post-hoc alignment mitigates this problem by mapping new-model queries into the old-model gallery space while preserving the new learned representation. However, because independently trained models can differ in fine-grained representation structure, the orthogonal alignment remains approximate, leaving a residual angular discrepancy between the old-model query and the aligned new-model query. We study whether spherical linear interpolation (SLERP) between these two normalized query representations can improve retrieval without re-indexing the gallery. We formalize the geometric conditions under which post-alignment interpolation yields a query direction closer to a task-optimal direction than either endpoint, and connect this characterization to Recall@$K$ through a local margin-based certification result. Experiments across multiple benchmarks and model families show that SLERP improves over orthogonal alignment alone, suggesting these conditions are broadly met in practice.
Chat is not available.
Successful Page Load