PAARBench: Evaluating Inference-Time Adaptation of Latent World Models
Abstract
Inference-time adaptation (ITA) of pretrained robot models is a strategy to handle dynamic environments, but it can also destabilize the closed loop, overfit on a small deployment history, or consume more compute than the gains justify. Binary task success alone obscures these failure modes. We introduce PAARBench (Plan--Act--Adapt--Repeat Bench), a reproducible benchmark for evaluating ITA strategies applied to visual latent world models under model-predictive control. PAARBench evaluates task success rates together with paired final-distance change, catastrophic degradation, compounding error, regret, adaptation latency, peak memory, and selection cost, each measured against a frozen reference. The initial release contains seven settings spanning in-distribution object pushing, held-out object shapes, a second visual pushing domain, and point-maze navigation under dynamics and layout shift. We benchmark six strategies: online gradient adaptation (OGA and OGA+), amortized hypernetwork adaptation (HyperLoRA), inverse-dynamics adaptation (PAD), a closed-form latent output correction (LEV), and an unconditioned low-rank control (StaticLoRA). The methods that accumulate an optimizer trajectory lead success across the push suite, yet each falls below its mode-matched frozen reference on at least one shifted maze setting. The methods that regenerate their correction from frozen weights before every replan are the most stable and give the largest gain on held-out maze layouts.