Can We Tell What GRPO Post-Training Added? A Cross-Scorer Audit Across Math and Code
Abstract
Any claim that post-training improved a pretrained model rests on a scorer that measures the difference. Does that conclusion depend on which scorer is used? We train Qwen-2.5-7B-Instruct with Group Relative Policy Optimization (GRPO) under five reward designs across math and code synthesis, holding the reported setup constant, and cross-score the resulting checkpoints. A reliability audit leaves all five math scorers usable, but invalidates the code-execution accuracy of every post-trained checkpoint, along with the three code checkpoints trained through that harness. On math the point estimates agree: all ten pairwise ranking correlations have tie-adjusted Kendall's tau_b >= 0.764, every scorer measures every trained checkpoint gaining 0.28 to 0.33 Pass@32, and the highest and lowest estimates of the same gain differ by at most 0.011. The surviving code evidence is the untrained base and three judge-trained checkpoints scored by two LLM judges. Both judges already score the base above 0.95, and the gains they report are 0.0024 to 0.0143. Within that narrow band, the weak judge scores two seeds of one training configuration at +0.0048 and +0.0143, whereas the strong judge scores both at +0.0071; the judges also give different checkpoint orderings. Per-prompt judge verdicts were not cached, so these comparisons have no uncertainty intervals. We also withdraw an earlier prompt-overlap argument. At n = k = 32 each per-prompt Pass@32 outcome is binary and flip probabilities vary by prompt, so allocating changes uniformly across prompts is not a justified null; matching directions are deterministic when two binary outcomes both differ from the same base value. In this model and task setting, the math conclusion is insensitive to the five scorers, while the current code evidence does not establish whether training helped. For pretrained models, where the final verifier and its eligibility mask do apply, the same judge-versus-execution gap appears on a second model family and does not close at a fourfold larger sampling budget. We recommend reporting the base score under every available scorer and treating near-saturation as a warning that uncertainty and independent evaluation are needed before certifying a gain.