Prediction Error Is the Wrong Trigger: Online Staleness Signals Dissociate from the Value of Adapting a World Model
Abstract
A robot that plans with a learned world model has to decide when the model is stale enough to be worth fine-tuning online. Fine-tuning costs compute and can make the planner worse, and the usual rule is to watch a prediction-error signal and fine-tune when it rises, which assumes that error signals track the value of adapting. We test that assumption on a stochastic pursuit-evasion task. A frozen model-predictive planner is deployed into twelve single-constant physics shifts, and for each shift we measure the adaptation headroom (the return gained by swapping in a model trained on the shifted physics) together with four online staleness signals computed from the executed trajectories at no extra environment cost: one-step likelihood, ensemble disagreement, planning-horizon rollout error, and the value gap between imagined and realized return. The signals do not track headroom. Headroom is concentrated in a single shift (+23.2 return), but one-step surprise fires just as strongly on shifts with no headroom, and the value gap fires hardest where a fresh model would not help while staying inside its calibration band on the one shift where adaptation pays. In closed-loop streams, the surprise gate almost never closes, so it behaves like always adapting, and the value-gap gate adapts intermittently and, on the one shift where adaptation pays, does worse than both extremes. A simple value-based trigger fixes this: run the adaptation as a short trial under interleaved control and commit only if the realized returns improve significantly. It refuses all 24 unfixable streams exactly, recovers the full gain of always adapting where it commits, and its only errors are conservative misses. Deciding when to adapt inherits the objective mismatch known from training world models: signals that measure how wrong the model is do not measure what fixing it is worth.