SPhyR: An Environment for Spatial-Physical Reasoning with a Simulation Verifier
Philipp Siedler
Abstract
Reinforcement learning from verifiable rewards needs verifiers, and the available ones are narrow: exact match, unit tests, or a learned reward model. Each assumes a task has a checkable canonical answer. Many design problems do not. We present SPhyR, an environment for spatial-physical reasoning whose verifier is a physics simulation rather than a comparison to a stored answer. An episode presents a 2D domain with prescribed loads and supports and a masked material distribution; the agent submits a completion, and the environment solves it as a linear elastic structure and scores it against a topology optimizer re-run on exactly the masked region. A completion that routes the load differently from the reference but just as efficiently scores just as well. Across the $24{,}325$ valid completions in our results, $72.7\%$ carry their load correctly while differing from the reference somewhere; a string-matching verifier rejects them all, and graph connectivity disagrees with compliance on $1{,}079$ of them, so a cheaper structural proxy will not substitute. We show the two views rank models differently and in both directions, that the environment's reward is not reachable by retrieval, and that a benchmark of this shape can rank models by grid-serialisation ability rather than by reasoning unless that is measured separately. The environment is released in the OpenEnv interface, single-step or with simulator feedback between attempts (Footnote: The dataset, environment and evaluation code are released; links are withheld for double-blind review and the code accompanies the submission as supplementary material.).
Chat is not available.
Successful Page Load