2nd Embodied Spatial Reasoning (ESR) Workshop
Abstract
Embodied spatial reasoning is the ability to analyze and interpret positions, orientations, physical properties, and temporal dynamics of an agent and its surrounding objects in 3D space. Rooted in core vision tasks like 3D detection and pose estimation, this capability serves as a foundation for advanced spatial understanding, reasoning, and generation in modern embodied AI and world modeling. However, prior discussions on this topic have been fundamentally fragmented. While embodied multimodal agents derive strong spatial reasoning capabilities from multimodal alignment, world models often extract powerful priors from the video domain. Yet these two lines of work lack a shared and reliable interface for spatial, physical, and temporal modeling. In particular, current systems often struggle to maintain object permanence, spatial memory, physical consistency, and long-horizon temporal coherence when agents move, interact with objects, or revisit previously observed regions. In this workshop, we invite prominent researchers across embodied AI, spatial reasoning, and world modeling to discuss the frontiers of this intersection. For example, we will discuss various spatial priors that enable agents to ``imagine'' and plan future actions, while endowing video world models with strict physical and temporal consistency. Furthermore, we encourage submissions that explore how agents actively learn from experience, leveraging reinforcement learning and/or curiosity to build action-conditioned spatial priors.