Spatial Representation Distillation and Knowledge Routing for Vision-Language-Action Models
Abstract
Understanding the 3D physical world has emerged as an essential component in vision-language-action models (VLAs), enabling robots to perform several challenging tasks by unifying \textit{spatial} perception, language, and action. Existing approaches typically either rely on explicit 3D inputs, such as depth maps or point clouds, or implicitly inject spatial knowledge into visual representations. However, explicit 3D inputs are often brittle in practice due to sensor noise, hardware heterogeneity, and incomplete depth coverage, while implicit alignment can distort the original 2D visual representation and weaken generic scene understanding. In this paper, we present SPARK-VLA, a simple yet effective framework to allow VLAs to implicitly possess generic 2D and spatial 3D visual comprehension within a unified visual backbone. SPARK-VLA consists of two key components: (1) spatial representation distillation, which transfers 3D understanding from a frozen 3D expert into a LoRA-adapted visual branch by aligning attention responses while preserving the original visual pathway, and (2) a lightweight knowledge router, which selectively forwards action-relevant visual tokens to the language model using routing supervision derived from action-loss gradients. Experiments in both simulation and real-world environments show that SPARK-VLA, with only a 0.5B language backbone and 50M trainable parameters, achieves performance comparable to or better than substantially larger VLAs.