Planning with a View World Model
Abstract
Embodied agents change not only the world through action, but also what part of the world they can observe. We study a view world model: an action-conditioned model of how an agent’s observations change under viewpoint motion, and whether that model supports deciding which future view to acquire. We operationalize this problem as view planning in VIEWSUITE, a 3D environment and benchmark built on real ScanNet scenes with full 6-DoF camera control. Path-to-View and View-to-Path probe whether a model can predict or invert known view transitions, while Interactive View Planning requires composing such transitions into a multi-turn plan that localizes an unseen target view. Across 13 frontier VLMs, the best models exceed 70% on short-horizon transition tasks but reach at most 21.3% on interactive planning, exposing a gap between local transition knowledge and action-conditioned future-view reasoning. To address sparse rewards, we aggregate every on-policy exploration trajectory, including failed ones, into a view graph. This graph is an empirical observation-space world model: its action-labeled paths are grounded future-view rollouts. Distilling reformulated graph paths back into the policy, interleaved with further self-exploration, improves Qwen2.5-VL-7B from 2.5% to 47.8% on interactive view planning and yields priors that transfer to other view-understanding tasks. Our results position future-view modeling as a complementary world-model problem for embodied decision-making, while revealing that robust view imagination without visually encountering the target remains open.