FloorSAV: Elucidating Spatial Audio-Visual Context with 2D Floormap for AV-LLMs
Abstract
While 3D spatial reasoning in dynamic egocentric environments is crucial for embodied intelligence, Audio-Visual Large Language Models (AV-LLMs) lack explicit mechanisms to process and internalize global geometry directly from raw sensory streams. Existing approaches either require costly fine-tuning or underutilize the model's cross-modal reasoning capacities. In this paper, we propose FloorSAV, a novel framework that explicitly grounds spatial audio-visual context by rendering a dynamic 2D floormap. By integrating 3D point clouds, camera trajectories, spatial audio cues, and semantically grounded object landmarks, we inject this floormap into the AV-LLM as a synchronized stream with an egocentric video. Assisted by the floormap-aware prompting, the model compounds its multimodal capabilities to jointly reason over visual, auditory, and geometric cues in a single inference. Evaluations on SAVVY-Bench show improved overall spatial reasoning across both open-source and proprietary AV-LLMs. We further introduce a dynamic relativity benchmark that evaluates spatial reasoning across the viewpoints of two moving agents, where FloorSAV demonstrates its effect. These findings highlight the importance of explicit global spatial representations for spatial audio-visual understanding in MLLMs.