ThinkWorld: World Imagination Reinforces Embodied Action Reasoning
Abstract
Existing Vision-Language-Action (VLA) models either predict actions end-to-end or introduce intermediate reasoning through textual or visual representations. Although such reasoning often improves planning and policy learning, it is typically optimized to encode high-level plans or scene structure rather than how actions transform the world. We introduce ThinkWorld, a framework that learns embodied action reasoning through reinforcement learning with world imagination. Given an observation and task instruction, ThinkWorld generates a Chain-of-Dynamics (CoD), a sequence of action thoughts describing how the world should evolve as the action unfolds. The CoD predicts the future world state and is reinforced according to how closely this prediction matches the observed transition, encouraging the learned reasoning to capture task-relevant dynamics. These world-grounded CoDs connect imagined action consequences to executable robot control. Experiments on simulated benchmarks and real-world manipulation tasks demonstrate the effectiveness of ThinkWorld over end-to-end and reasoning-based VLA baselines.