When Does Latent Prediction Error Support Physical Surprise Claims in Video World Models?
Abstract
Prediction errors from video world models are often treated as physical surprise, but they can also reflect appearance, static content, motion, or ordinary prediction difficulty. We present an evaluation framework that separates what frozen features contain, how they organize appearance and physical differences, and what changes a model's own prediction error. We study IntPhys2, a synthetic benchmark of possible and impossible events, and exact appearance and motion changes rendered with Kubric, a controllable synthetic video generator. We compare frozen features from a self-supervised latent predictor, a masked pixel reconstruction model, and simple color and motion statistics. Appearance remains easy to recover across these features, while one evaluated self-supervised video model exposes only modest information about whether events are possible or impossible. Looking at the latent geometry, appearance changes become smaller, line up more consistently across scenes, and overlap much less with physical changes in deeper layers. Across three evaluated action-conditioned world models, incorrect actions consistently raise native prediction error, but appearance changes with state held fixed produce either larger or smaller errors than approximately matched opposite actions depending on the model and score. Exact Kubric rerenders produce similar error changes for visual and physical interventions. A model can therefore represent physical information and react to incorrect actions without guaranteeing that its native prediction error measures physical surprise.