Visualising Latent Rollouts in LeWorldModel
Abstract
Reconstruction-free world models such as LeWorldModel plan directly from latent rollouts, avoiding the computational cost of predicting future pixels. However, this efficiency removes a valuable debugging interface: when the planner evaluates candidate action sequences, the corresponding latent trajectories provide little insight into what the model expects to move. We study on-demand reconstruction as an interpretability layer for frozen LeWorldModel models, enabling candidate futures to be visualized without modifying the planner. Our experiments reveal that full-frame diffusion reconstruction is poorly matched to this objective: the model expends capacity redrawing static geometry and frequently hallucinates duplicated objects. We instead preserve static scene content from the current observation and reconstruct only the future trajectory of moving objects, conditioned on the latent rollout and candidate actions. Across TwoRooms, Reacher, Push-T, and OGBench Cube, this approach produces recognizable, temporally aligned visualizations while retaining scene structure. Push-T achieves a 97.4\% validation hit rate and 2.34-pixel rendered-centroid error, while TwoRooms achieves a 98.0\% validation hit rate with nearly perfect wall IoU. Rather than replacing the planner, on-demand reconstruction provides a visual probe for inspecting and comparing candidate latent futures before the controller executes the next action.