LeVLWM: Text-Goal Control with a Single Vision-Language World Model Download PDF
Qiuyi Hua ⋅ Ahitagni Das ⋅ Vivek Boominathan ⋅ Randall Balestriero
Abstract
Latent world models plan toward a goal image: to specify a task, one must already hold a picture of its outcome. Simulators render goal frames on demand, which makes this invisible in benchmarks and unavailable almost everywhere else. We ask what it takes to make such a latent actionable from language instead, and find the obstacle is not language understanding but well-posedness. Regressing a caption onto a single goal latent is sound only when the caption is an information-complete encoding of the goal state---attainable when the state is low-dimensional, and the binding constraint as scenes grow cluttered. Otherwise the regression target is a conditional mean over every state consistent with the caption, and the planner minimizes distance to a latent that no reachable state occupies. We introduce LeVLWM, the first fully pretrained single vision--language world model that meets this condition by adding to a latent world model a from-scratch text encoder and two cross-modal predictors, trained end-to-end under one summed objective with SIGReg as the only collapse control. Existing methods obtain language conditioning in a separate stage, relying on pretrained vision-language backbones, frozen text encoders, or post-hoc alignment; LeVLWM learns it jointly with the dynamics in a single end-to-end pretraining run---no reconstruction, no reward, no pretrained language model. A caption becomes a goal embedding that standard MPC plans toward unchanged. Across four 2D and 3D control tasks the same checkpoint succeeds on $82.3\pm8.5\%$ of episodes from language goals and $82.7\pm8.8\%$ from goal images ($5$ seeds $\times$ $50$ episodes per environment) under an identical planner and harness, a difference of $0.4$ points, and linear and MLP probes recover physical quantities from the text-goal latent as accurately as from the image latent: the world model cannot tell whether its goal came from a camera or a sentence. This identifies information-completeness, rather than fluency, as the property that makes a goal specifiable.
Chat is not available.
Successful Page Load