Probing Where Surprise Happens: Localized Violation-of-Expectation Evaluation of Video World Models
Abstract
Violation-of-expectation (VOE) benchmarks assess intuitive physics in video world models by testing whether physically impossible events elicit greater prediction error than matched possible events. However, aggregating errors across space and time leaves unclear whether successful discrimination coincides with sensitivity to the observable violation evidence. We introduce localized VOE evaluation, a framework that uses spatiotemporal violation evidence annotations to measure possible--impossible prediction error differences within violation event relevant regions and intervals. We compare global, temporally localized, and spatiotemporally localized evaluations of V-JEPA models pretrained on developmental egocentric and large-scale natural-video datasets using Physical Concepts and IntPhys1. Global and localized judgments can diverge in both directions: aggregation can obscure correct localized discrimination or yield correct global discrimination despite reversed sensitivity within the violation region. Across BabyView pretraining, global--local agreement evolves non-monotonically, while local judgment confidence increases during continuous training despite modest accuracy gains. These findings establish localized VOE as a complementary diagnostic and show that aggregate discrimination alone does not establish sensitivity at the location and time of observable physical violations.