PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning
Qiran Zhang ⋅ Yuheng Wang ⋅ Runde Yang ⋅ Lin Wu ⋅ Jingru Fan ⋅ Shu Yao ⋅ Jie Zhang ⋅ Tianle Zhou ⋅ Huatao Li ⋅ Ruijie Shi ⋅ Yihan Li ⋅ Chen Qian
Abstract
Programmatic video generation through code offers geometric precision and temporal coherence unattainable by pixel-level diffusion models, yet rigorously evaluating whether language models can produce spatially correct animated output remains an open problem. We introduce **PRISM**, a large-scale benchmark of 10,372 human-calibrated instruction-code pairs ($20\times$ larger than prior Manim benchmarks) grounded in real-world educational scenarios across English and Chinese, spanning 437 subject categories. Alongside the dataset, we propose a funnel-style evaluation framework with four complementary metrics: *Code-Level Reliability* as the executability gate, *Spatial Reasoning* as the core end-to-end metric measuring layout correctness over full animation sequences, and *Prompt-Aware Dynamic Visual Complexity* (PADVC) and *Temporal Density* (TD) as diagnostic dimensions characterizing dynamic expression and temporal activity. Systematic evaluation of seven mainstream LLMs reveals a striking *Execution-Spatial Gap*: the average drop from execution success rate to spatial pass rate is approximately 41\%, demonstrating that syntactically correct, runnable code does not automatically yield spatially coherent visual output. No single model dominates all dimensions; notably, the model with the highest execution rate exhibits the largest Execution-Spatial Gap, while the most spatially robust model does not lead on code reliability. These findings highlight that evaluating programmatic video generation should not stop at executability. PRISM provides a principled benchmark for advancing spatially coherent code generation.
Chat is not available.
Successful Page Load