PhysEval: Quantifying the Gap Between Video Generation and World Physical Laws
Abstract
Text-to-video (T2V) models can now synthesize visually compelling motion, but visual plausibility does not guarantee that the generated world follows the physical laws it depicts. A falling object, sliding block, collision, or oscillating spring may look reasonable while still implying an incorrect acceleration, material parameter, or conservation relationship. Most existing benchmarks emphasize perceptual quality, text-video alignment, or qualitative physical plausibility, leaving a gap between visually plausible video generation and quantitatively verifiable physical behavior. To quantify this gap, we present PhysEval, a benchmark dataset and automatic evaluation protocol for scoring the physical accuracy of T2V generations. PhysEval pairs each prompt with auditable metadata, including the target value, unit, evaluator type, expected object count, calibration setting, and known physical parameters. It covers ten physical metrics across kinematics, physical constants, material parameters, and conservation laws. Given generated videos, PhysEval automatically determines which samples are measurable, estimates the relevant physical quantities, and produces normalized scores together with discard diagnostics. Our evaluation and detailed analysis show that current T2V models still struggle to reliably generate videos that are both measurable and physically consistent: even visually plausible samples often fail to express the quantitative cues needed by the underlying physical law.