How Do Video Foundation Models Encode Intuitive Physics? Probing Across Pretraining Paradigms
Abstract
Do video foundation models encode a usable sense of how the physical world works, and does intuitive physics emerge naturally from large-scale video pretraining? In this work, we study whether pretrained video foundation models encode intuitive-physics information in their frozen representations, and how its accessibility varies across pretraining paradigms, layers, backbone variants and probe types. Using frozen-feature probing on IntPhys2 and Minimal Video Pairs (MVP), we compare predictive joint-embedding models (V-JEPA), masked reconstruction models (VideoMAE), and diffusion-based video generators (LTX-Video). We find that across linear, MLP, and temporal attentive probes, physics-relevant information is generally weakest in early layers and most accessible at intermediate to late depth. On MVP, temporal attentive probes substantially outperform linear readouts, with V-JEPA achieving 95.0% pair consistency. On IntPhys2, linear probes recover more signal, and VideoMAE attains the strongest temporal attentive result at 73.9% violation-of-expectation accuracy. On both benchmarks, attentive probe performance on LTX consistently underperforms compared to the other two pretraining paradigms. Together, these results suggest that intuitive-physics is broadly decodable from video foundation models, but it is not represented as a single, uniform capability. Its accessibility—and especially its temporal grounding—depends much more on the benchmark and readout than on a simple hierarchy of pretraining objectives or model scale.