Impossible Code, Barely Changed Score: A Replay Audit of Judge-Based Evaluation
Abstract
Training a language model with reinforcement learning requires a scorer that decides whether an answer is right. For code the dependable option is to run the program against tests, but that is slow and needs a sandbox, so many studies ask a large language model to judge the code by reading it instead. This paper reports what that substitution cost inside one real evaluation. A finished study had trained Qwen-2.5-7B-Instruct with Group Relative Policy Optimization (GRPO) and saved the programs its checkpoints wrote for 421 coding problems. We later downloaded the same checkpoints and asked the same questions again. The untrained model answered as before, failing to compile on 0.40% of attempts against 0.39% the first time. The two trained checkpoints came back writing code that will not compile at all, on 63.0% and 66.9% of attempts, almost always through a single surplus closing bracket. The judge barely reacted: scoring the same checkpoint in both runs at a matched budget, its reported accuracy fell from 0.9667 to 0.9395, while the share of its outputs that were valid Python fell from 99.6% to 33.1%. Executing the programs shows how far apart judgment and behaviour had drifted: its Pass@32 is equivalent to about 99 of the 421 problems solved, against about 396 under the judge. The judge is not blind, but how well it separates the damaged checkpoint from the base depends on how many attempts each problem is allowed: 0.1601 at one attempt, 0.0215 at 32 and 0.0071 at 128, while execution holds between 0.2095 and 0.1496 over the same range. We propose one free check beside every judge-based number: whether the output parses at all. It costs a healthy run 0.0019 accuracy and a damaged one 0.4051, and of 70,239 non-parsing samples execution credited none.