Beyond Video Generation: Exploring Latent Interaction Priors from Frozen Video World Models for Robot Learning
Abstract
Pretrained video world models learn conditional representations of visual scene evolution, but are often accessed through future video generation, multi-step inference, or robot-specific adaptation. We explore an alternative interface that reuses the initial conditional response before future synthesis. Given an initial observation and task instruction, we extract this response from a frozen Flow Matching video model at the noise endpoint and expose it as a Latent Interaction Prior (LIP) for downstream manipulation. A lightweight Perceiver encodes it into compact prior tokens, while the video backbone remains frozen and no future frames are synthesized. With Diffusion Policy, DP + LIP achieves 90.6% average success on LIBERO and 65.7% and 22.4% success rates on the Easy and Hard settings of RoboTwin 2.0, respectively. Further analysis provides evidence that LIP supplies task-conditioned latent information beyond static features from the same video backbone. The interface also transfers to a jointly trained pi0.5 policy on RoboTwin 2.0 and a pi0 policy validated on the AgileX PiPER platform. These results suggest that the initial response of a pretrained video model can serve as a compact latent prior for robot manipulation, providing an alternative to explicit future video generation.