When Does RLVR Teach Physical Strategy? Measuring Simulator-Only Post-Training for CNC Planning
Shishun Dai
Abstract
When does reinforcement learning from verifiable rewards (RLVR) give a language model a new physical strategy, and when does it only improve program executability? We measure that split on CNC process planning, where actions must respect stock geometry, tool contact, and machine constraints. An open model writes complete multi-setup G-code. An interpreter and voxel process model grade each proposal and supply the only post-SFT training signal. The grader passes 166 checks and matches an independent 2.5D engine on applicable cases. The teacher never uses multi-tool roughing and finishing or reorders nonadjacent same-face features into one setup. Both behaviors are described in the prompt and rewarded by the grader. SFT reaches $86$--$89\%$ path validity but only $15$--$18\%$ strict success in-distribution and $0/100$ on both exams. Simulator-only GRPO improves executability (path validity $86.0\%\rightarrow93.0\%$ IID and $59.0\%\rightarrow76.0\%$ on the setup exam) and mean verifier reward ($-0.417\rightarrow-0.344$ validation). Strict success changes by only $0.5$ percentage points on validation and IID, while both exams remain at $0/100$ with $0\%$ on their strategy proxies. Paired validation outputs contain two SFT-failure to GRPO-success repairs and one regression; IID contains three repairs and two regressions. Thus, individual repairs do not constitute an aggregate strategy gain. We do not claim a larger budget cannot close the gap. In this single-seed regime, post-training improves executability without recovering either tested strategy. We provide the audited grader, frozen benchmark, and archived outputs for offline regrading in an anonymized bundle.
Chat is not available.
Successful Page Load