Benchmarking Fine-Grained Spatio-Temporal Awareness in Embodied Brain Models
Abstract
Large Vision-Language Models (LVLMs) have shown remarkable potential for embodied AI, yet a critical bottleneck persists: the lack of fine-grained spatio-temporal alignment between high-level textual reasoning and low-level visual grounding. Existing benchmarks fail to recognize the critical importance of such fine-grained alignment for embodied tasks. To address this gap, we introduce ECLBench, a comprehensive benchmark designed to systematically assess spatio-temporal embodied understanding across two primary capabilities: Embodied Cognition (Object Cognition and Spatial Cognition) and Embodied Localization (Grounding and Pointing). Together, these four pillars encompass 21 specialized sub-capabilities, supported by 3,616 egocentric video clips (577,998 frames) and 12,000 meticulously curated open-ended questions. Extensive evaluations on ECLBench reveal that general LVLMs excel in semantic understanding but struggle with precise spatio-temporal grounding, while existing embodied agents often sacrifice generalization for specialization.