Fooling One Fools Them All: Universal Over-Optimization in Model-Graded RL
Abstract
Reinforcement learning has been crucial for helping improve models in domains with cheap verifiable rewards such as coding. However, many useful tasks, such as alignment research and conceptual thinking, lack such rewards. Moreover, reliable human feedback is expensive to collect for such tasks. Model graders therefore offer a scalable alternative, and can serve as both a reward signal during training and as a proxy for tracking task progress. However, policies can learn to over-optimize against these graders during training, so stronger held-out graders are often used to track progress more reliably. Our work questions this assumed reliability, demonstrating the existence of universal over-optimization, where scores from training and held-out graders, across model families and scales, increase without a corresponding improvement in task performance. Notably, in a chess environment, the trained policy produced incorrect responses that received high rewards from all held-out graders that we tested. Similarly, in a creative writing environment, all graders preferred writing stories from the trained policy, while human evaluators preferred stories from the untrained policy. These results suggest caution when relying heavily on model-graded evaluations, especially those involving sandwiching techniques to study and improve scalable oversight.