PhysAgentGym: Free Physical-Law Verifiers for Training Small Code-Reasoning Agents
Abstract
Code-reasoning agents trained from rollouts typically rely on costly verifiers — expert annotation, hand-written unit tests, or learned reward models — each of which scales poorly across domains. We argue that in domains with closed-form physical laws, the laws themselves form a free, dense, automatic class of reward signals. We instantiate this for physics via PhysAgentGym, an agentic code- execution gym in which a language model receives a physics scenario, writes Python that simulates it, and emits a trajectory scored by a rule-based verifier that compiles natural-language standards into executable trajectory predicates. Verifier signals are label-free: physical laws supply the ground truth without any per-task annotation. Across four physics subcategories totaling 1,320 closed-form- simulator task instances, two frontier models, Claude Sonnet 4.5 and OpenAI GPT-5, achieve nearly-identical aggregate scores (0.944 and 0.945 at N=30 each), with per-cell agreement within 0.5 pp, providing independent evidence the verifier is well-calibrated rather than arbitrary. We filter 491 successful frontier trajectories on two in-distribution subcategories into an SFT corpus and train Qwen-2.5-Coder {1.5B, 3B, 7B} with QLoRA. The trained 7B adapter reaches mean = 0.942 on the four-subcategory test set, matching both frontier models within 0.3 pp on aggregate at approximately 50×lower inference cost. It scores identically to both Sonnet and GPT-5 on all three in-distribution subcategories and ties them within 0.8 pp on out-of-distribution Damping, despite never seeing damping data in training. We further show that combined sparse-and-dense reward filtering beats sparse-only filtering by 1.4–8.7 pp, that smaller models exhibit fundamental seed brittleness which disappears at 3B and above, and that scale enables the OOD generalization that 1.5B and 3B do not provide. The method is portable to any domain admitting closed-form physical laws