Mimicking the Physicist's Eye: A VLM-centric Approach for Physics Formula Discovery
Jiaqi Liu ⋅ Songning Lai ⋅ Pengze Li ⋅ Di Yu ⋅ Zhou wenjie ⋅ Yiyang Zhou ⋅ Peng Xia ⋅ Zijun Wang ⋅ Xi Chen ⋅ SHIXIANG TANG ⋅ LEI BAI ⋅ Wanli Ouyang ⋅ Mingyu Ding ⋅ Huaxiu Yao ⋅ Aoran Wang
Abstract
Automated discovery of governing equations remains challenging when models rely only on numerical tokens and ignore the phase-space and trajectory evidence that is informative in low-dimensional dynamics. We propose Visual Induction for Physics-based Equation Reasoning (VIPER-R1), a multimodal framework for visual formula discovery from phase portraits, trajectory plots, and aligned measurements. VIPER-R1 uses a supervised Motion Structure Induction (MSI) curriculum for joint rationale-and-equation generation, and is then refined by \textbf{Reward-Guided Symbolic Calibration (RGSC)}, a GRPO-based stage optimized with a structure-aware reward that emphasizes topological correctness over coefficient matching. At inference time, we optionally apply Symbolic Residual Realignment (SR$^2$), an external symbolic regression module that models only the residual numerical mismatch of the base hypothesis. We release PhysSymbol, a 10{,}000-instance two-part multimodal corpus; the quantitative evaluation in this paper uses its Part 1 mechanics split, which provides paired visualizations, trajectory data, and expert C-CoT traces. On this benchmark, VIPER-R1-7B achieves the best structural score ($S_{\mathrm{struct}}=0.812$), is effectively tied with VIPER-R1-3B on exact-match accuracy ($S_{\mathrm{acc}}=0.487$ vs.\ $0.488$), and, after optional SR$^2$ refinement, reaches a Post-SR$^2$ MSE of $0.032$, nearly $3\times$ lower than the best VLM baseline. These results show that VIPER-R1 improves symbolic structure induction in focused low-dimensional dynamics by combining paired visual evidence with structure-aware calibration to produce stronger first-order hypotheses for downstream refinement.
Chat is not available.
Successful Page Load