Assessing the Gameability of Vision Language Model Judges for World Model Video Evaluation
Abstract
Vision-language model (VLM) judges are increasingly used to evaluate the physical plausibility of world model videos and, in some settings, as optimization signals for video generators. We stress test three specialized judges: WorldModelBench VILA, VideoPhy-2-AutoEval, and PhyJudge-9B, using paired clean and perturbed videos designed to test sensitivity to physics-relevant temporal corruption and invariance to physics-preserving superficial cues. Temporal effects are evaluated against blinded human judgments, while superficial perturbations are compared against a frozen V-JEPA representation with a human-supervised ordinal plausibility probe. In a blinded human audit, temporal shuffle, freeze, and reverse reduced Physical Commonsense scores by 0.89-2.66 points on a 1-5 scale, while the judges showed much smaller declines or slight increases. Conversely, every text overlay family increased scores for at least two judges despite negligible changes in the reference plausibility score. A contentless blank box of matched size and opacity produces a similar observed mean effect on one judge to its rubric-vocabulary overlay, so the shortcut is not always linguistic, and a small search over overlay wording raises judge reward further and transfers across judges. These perturbations can also alter generator rankings. Together, our results show that specialized world model video judges can under-respond to physically meaningful temporal corruption while remaining sensitive to superficial visual cues, raising concerns about their reliability for both benchmarking and optimization.