SphMind: Towards Robust, Training-Free VLM-based Spatial Reasoning with a 360 Camera
Abstract
Omnidirectional or 360° cameras provide embodied AI agents with a holistic, wide field-of-view (FoV) perception of the 3D structure of their surroundings. This capability has generated significant interest in applying Multi-modal Large Language Models (MLLMs) to omnidirectional spatial reasoning. However, most MLLMs are primarily trained on conventional 2D perspective images and therefore struggle with the severe distortions and wrap-around discontinuities introduced by spherical geometry. As a result, enabling MLLMs to generalize effectively to non-Euclidean 3D spaces without retraining remains an open challenge. In this paper, we propose SphMind, a novel, training-free, and plug-and-play framework designed to bridge this gap. The central idea of SphMind is to decouple semantic perception from geometric reasoning. Instead of requiring MLLMs to learn complex spherical geometric principles internally, the framework preserves their strong semantic understanding while handling geometric reasoning externally. To achieve this, we introduce a Spherical Harmonics-based Spatial Graph (SHSG) that models spatial relationships using equivariant transformations on the sphere. We further integrate this with Inference-Time Geometric Grounding (IGG), a model-agnostic closed-loop optimization process that aligns the internal representations of MLLMs with spherical geometric constraints during inference. Extensive experiments on three benchmark datasets demonstrate the effectiveness of SphMind. Without any additional training, the framework achieves more than 21.4% average improvement in directional reasoning on MP3D and Stanford2D-3D, outperforms all prompt-engineering baselines by 8.7% on the real-world ODI-Bench dataset, and achieves 5.9× higher rotational invariance under panorama rotations compared to existing baselines, all without dataset-specific tuning. We also validate the framework on real-world in-the-wild captures, where SphMind successfully resolves directional reasoning queries that baseline vision-language models fail to answer correctly.