PhysWorldBench: Benchmarking Video Generators as Physical World Simulators
Abstract
While video generation models demonstrate impressive visual fidelity, their capacity to serve as reliable physical world simulators remains an open question. Existing physical benchmarks face two structural limitations: Vision-Language-Model (VLM) judges hallucinate on fine-grained physics, and pixel-level comparisons conflate visual deviation with physical violation, neither localizing which law was broken. To bridge this gap, we introduce PhysWorldBench, a faithful benchmark for physical video generation covering four classical macroscopic domains observable from video: Rigid Body Dynamics, Optics, Acoustics, and Fluid Mechanics. Our framework pioneers a tool-augmented evaluation pipeline: for each scenario, we decompose the physics into fundamental governing equations and conservation laws, extract the variables that appear in those equations directly from videos using deterministic computer-vision and signal-processing tools, and quantify adherence as a per-law residual against simulator ground truth or analytic invariants. This grounds the evaluation in objective measurements rather than judge ratings, eliminates VLM-induced biases, and pinpoints which physical law a model violates rather than only that a visual deviation occurred. We instantiate the benchmark across 241 phenomena and 5,522 generated videos from 13 video generators. The results quantify a substantial visual-physical gap: state-of-the-art models render plausible textures yet routinely violate energy, mass, and geometric-optics constraints, providing a robust evaluation suite for advancing physically grounded video world models.