WOVEN: Weaving Visual World Modeling into Multimodal LLMs
Abstract
Multimodal large language models (MLLMs) fail across a striking range of embodied and robotic reasoning tasks: embodied action understanding, physical dynamics, spatial mental modeling, and temporal benchmarks each report near-chance accuracy on their hardest categories, and each failure is documented and attributed separately. We hypothesize that these failures share one primitive: action-conditioned visual transition reasoning, the ability to tie an action to its visual consequences over a state transition (s, a, s'), which forms the core of visual world modeling. To operationalize this hypothesis, we harness a video generative model (VGM) as the data generator, exploiting the controllable, diverse, and realistic rollout prior that large-scale next-frame pretraining confers. On this engine we build WOVEN, 36,076 four-way questions that weave 5 action types, 8 reasoning types, and 20 scene types into a cognitive-theory-guided, cross-factored grid with densely typed distractors. Evaluating 38 frontier MLLMs confirms the deficit: transitions produced by a 14B VGM confuse models many times its size, the best model trails the human ceiling by roughly 30 points, and the deficit resists scale. We ask whether visual world modeling can be taught as a shared primitive. Post-training on WOVEN transfers zero-shot to 22 of 26 external benchmarks with gains up to +27 points (embodied action understanding +11, robot manipulation QA +9), persists from 3B to 32B, and WOVEN items substitute for half of a benchmark's own training data at no cost: the taxonomy trains a capability that downstream tasks genuinely share. Finally, we ask what good world-model training looks like for MLLMs, and extract a predictive, actionable recipe from controlled contrasts across the taxonomy: the reasoning type of supervision, unlike its action or scene type, sets where transfer lands; its intervention locus sets which perturbations the model withstands; causal supervision erases the default-state prior; and temporal supervision alone carries temporal localization. These results establish visual transition reasoning as a measurable, trainable, and broadly transferable core of world modeling, and lay the groundwork for future work to diagnose emerging MLLMs and compose better world-model training recipes.