Solver to System: LLM Reasoning Across Representations and Execution in Scientific Simulation
Abstract
Simulators based on partial differential equations (PDEs) are widely used to model physical systems, but generating and interpreting PDE solver code with large language models (LLMs) remains challenging. Scientific simulation code is a useful testbed for LLM scientific reasoning because (1) one system can be represented by code, equations, natural language, and numerical trajectories and (2) code can execute successfully while implementing invalid physics. We ask whether LLMs can recover the underlying physical system across representations and under lexical perturbations, evaluating (1) cross-representation consistency and (2) static and agentic detection of physical invalidities in code. We introduce Solver2System, a dataset of PDE solvers with multiple representations, controlled corruptions, and lexical perturbations. Across eight open-weight models, LLMs systematically favor some representations when evidence conflicts, and lexical perturbations to code alter interpretation of other unchanged representations. Execution improves physical-validity judgments; weaker agents can judge correctly without identifying causal defects in code, while stronger agents combine accurate static judgments with targeted diagnostics. Together, these results show that verification accuracy can overstate LLM scientific reliability and reasoning: LLMs can reach correct judgments with inconsistent representation reconciliation or causal misdiagnosis.