Planning with View World Models
Abstract
A world model for physical AI must represent not only how a scene changes, but also how the agent's accessible evidence changes under its own motion. We isolate this observation-dynamics problem in static real 3D scenes. A view world model predicts or organizes action-conditioned future observations under camera movement, and its utility is measured by whether an agent can choose future views that support a goal. ViewSuite provides full 6-DoF camera control over reconstructed ScanNet scenes and separates local forward and inverse transition reasoning from multi-turn planning. Across 13 frontier VLMs, the strongest models exceed 70% on short-horizon Path-to-View or View-to-Path tasks, yet interactive view-planning success reaches at most 21.3%. Increasing the interaction budget yields limited gains, and replacing point-cloud rendering with higher-fidelity 3D Gaussian Splatting changes interactive performance by only 0.2 to 1.9 percentage points for the evaluated models. Visual fidelity alone therefore does not resolve the action-composition bottleneck. We construct a discrete observation-space world model by aggregating every on-policy transition into an action-labeled view graph. Its paths provide experienced future-view rollouts that can be reformulated as goal-conditioned supervision. Alternating graph distillation with self-exploration raises Qwen2.5-VL-7B from 2.5% to 47.8% interactive success. The results motivate downstream planning utility as a complement to pixel fidelity when evaluating world models, while leaving generalization to unvisited views, dynamic scenes, and physical interaction open.