Predictive Sensitivity Across Contact Transitions in a Video World Model
Abstract
Video world models are increasingly used as learned simulators for robot manipulation, but a model that predicts its own rollout well is not necessarily a model that responds correctly to small changes in the state it is predicting from. We study this distinction around physical contact transitions in iVideoGPT, an autoregressive video world model, using a simulator-grounded causal intervention: we restore an exact simulator state, apply a small, calibrated physical perturbation to the manipulated object, and compare the two resulting model rollouts. Natural rollout error shows a distinguishable signature at contact release, but not onset, after controlling for a ceiling confound. A matched causal test at a fixed offset before each transition finds the same asymmetry causally: the release-approach effect independently replicates and survives restriction to transition-isolated windows, while the analogous onset effect fails independent replication and is discarded. A final event-centered experiment sweeps the intervention across seven offsets and finds that local sensitivity does not jump at either transition; it rises smoothly through onset and falls smoothly through release, with within-event contrasts showing the same trend when the event set is held fixed across offsets. The single offset tested in the matched experiment sits on the low side of onset's rising profile and the high side of release's falling profile, one reason the two transitions looked so different at a single point. A qualitative and tracker-based audit shows the underlying pixel and object-position differences are real, object-localized, and graded with the causal-sensitivity result, but individually small, often below the resolution of a validated cube tracker, so we report this only in aggregate. Predictive sensitivity around physical interaction is temporally structured and directionally asymmetric between contact onset and release, a distinction plain rollout-error metrics do not by themselves reveal.