Multi-view Relational Distillation for Spatial Reasoning with Vision-Language Models
Abstract
Vision-language models (VLMs) have achieved strong image and video understanding, yet their visual-spatial representations remain geometrically fragile, leading to failures in spatial reasoning required for embodied AI, robotics, and autonomous driving. Existing approaches to geometry grounding either fine-tune VLMs on spatial question answering, which can perpetuate spurious visual representations, or fuse features from large geometry-grounded vision models, which substantially increases model size during inference. Knowledge distillation from geometry-grounded vision models offers an alternative, but directly matching multi-view teacher features can disrupt the pretrained alignment between visual and textual representations, degrading object- and language-semantic understanding. We propose Multi-view relational distillation (MVRD), which distills patch-wise cosine similarities across views instead of the teacher features themselves. These multi-view relations encode geometric correspondences sufficient for spatial understanding while leaving the student representation underdetermined, allowing it to remain close to its pretrained vision-language space. Across representative VLMs, MVRD improves visual-spatial reasoning, outperforming supervised fine-tuning and direct feature distillation while approaching feature-fusion methods with substantially fewer added parameters and lower latency. Further analysis shows that MVRD makes visual representations more geometric without breaking vision-language alignment, and generalizes to 3D scene understanding tasks, including object grounding, dense captioning, and question answering.