WM-LAN: Test-Time Adaptation of World Models through Latent Refinement
Ankit Bhattarai ⋅ Matthew Macfarlane ⋅ Clem Bonnet ⋅ Florian Fischer ⋅ Per O Kristensson
Abstract
General agents that act in the real world need to adapt to environments they have never seen before. World models are a promising architecture for the core of a general agent, which it can use to plan, predict the consequences of its actions, and learn. However, world models are only useful when they are accurate. When entering unseen environments, agents must continually adapt their world models on the fly in a sample-efficient manner. Amortised inference produces a context estimate in a single forward pass, but the estimate is fixed once training ends and degrades out-of-distribution (OOD), with no way to spend more compute to improve it. We propose the World-Model Latent Adaptation Network (WM-LAN), which adapts a world model at test time through latent refinement. WM-LAN uses an encoder--decoder architecture and adapts by searching the latent space for the representation that, when decoded, best explains all transitions seen so far. The search requires no reward, labels, or parameter updates. On CartPole and Walker with unseen gravity and actuator strength, we show that amortisation error accounts for most of the OOD error of an encoder-decoder world model. Refinement from 32 transitions removes $51$-$76$\% of the OOD one-step error, and with more observations it matches a decoder conditioned on the ground-truth context. On CartPole, the benefit of refinement grows with planning horizon: compounding error in the amortised model makes long-horizon planning worse than not planning at all, whereas planning through the refined model yields the largest gains over the policy acting alone.
Chat is not available.
Successful Page Load