Benchmarking Fine-Grained Spatio-Temporal Awareness in Embodied Brain Models
Abstract
Large Vision-Language Models (LVLMs) have shown remarkable potential for embodied AI, yet a critical bottleneck persists: the lack of fine-grained spatio-temporal alignment between high-level textual reasoning and low-level visual grounding. Existing benchmarks fail to recognize the critical importance of such fine-grained alignment for embodied tasks. To address this gap, we introduce ECLBench, a comprehensive benchmark designed to systematically assess spatio-temporal embodied understanding across two primary capabilities: Embodied Cognition (Object Cognition and Spatial Cognition) and Embodied Localization (Grounding and Pointing). Together, these four pillars encompass 21 specialized sub-capabilities, supported by 3,616 egocentric video clips (577,998 frames) and 12,000 meticulously curated open-ended questions. Extensive evaluations on ECLBench reveal that general LVLMs excel in semantic understanding but struggle with precise spatio-temporal grounding, while existing embodied agents often sacrifice generalization for specialization. Guided by the benchmark's findings, we develop ECLBrain, a baseline model that discretizes continuous spatial coordinates into integer tokens. This simple yet effective design bridges textual semantics and spatio-temporal vision, demonstrating a promising path toward unified embodied reasoning. ECLBench and ECLBrain together establish a rigorous, multi-dimensional standard for diagnosing current limitations and guiding future research toward spatially-aware and temporally-consistent embodied intelligence.