JEPA-TTT: Persistent Test-Time Training of Latent World Models for Planning under Dynamics Shifts
Abstract
World models enable agents to plan by predicting future states of the environment, but their predictions can become unreliable when test-time dynamics differ from those seen during training. We introduce a persistent test-time training method for pretrained action-conditioned Joint-Embedding Predictive Architecture world models, JEPA-TTT. It updates only the latent dynamics predictor using self-supervised prediction targets from a frozen visual encoder, and requires no goal image or online environment reward for planning. The adapted predictor is retained across episodes, allowing adaptation to accumulate over time. JEPA-TTT uses dense replay, which forms prediction windows at every temporal offset, retains them in a growing buffer, and samples minibatches from that buffer for predictor updates. Across eight dynamics shifts in four continuous-control environments, JEPA-TTT improves planning on every shift. After 500 test-time episodes, it reduces autoregressive latent prediction error by 83% on average and improves planning performance by 153% over the frozen world model. These results show that persistent self-supervised test-time training can adapt a pretrained latent world model under changed dynamics.