The Heel of RLVR: Benchmark Glory Should Not Outpace Honest Measurement
Shuo Yang ⋅ Chiyu Ma ⋅ Kexin Huang ⋅ Jinda Lu ⋅ Shaohang Wei ⋅ Xinpeng Liu ⋅ Haoming Meng ⋅ Yuyang Liu ⋅ Shangshang Wang ⋅ Minghao Zhu ⋅ Soroush Vosoughi ⋅ Guoyin Wang ⋅ Jingren Zhou ⋅ Li Yuan
Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as the dominant paradigm for eliciting reasoning capabilities in large language models, producing a steady stream of benchmark improvements that have generated considerable excitement in the community. In this position paper, we argue that this reported progress is built on a measurement foundation with systematic, largely unexamined cracks. Through large-scale quantitative analysis of rollout trajectories during Zero-RL training, we identify five structural vulnerabilities organized across two levels. At the algorithm level, we show that destructive over-reflection ("Oops Moments") occurs at nearly three times the rate of the celebrated self-correction ("Aha Moments"), exposing a profound survivorship bias in how progress is reported; and that injecting $\pm20$% noise into the advantage signal produces no measurable effect on training dynamics, calling into question what our optimization algorithms are actually learning. At the implementation level, we show that rule-based verifiers maintain a persistent misjudgment rate of $0.1$%--$0.5$% that silently corrupts reward signals throughout training; that output format choices orthogonal to reasoning ability measurably shift benchmark performance; and that the alignment between loss normalization granularity and micro-batch construction strategy constitutes a hidden hyperparameter that can cause two teams running ostensibly the same algorithm to optimize fundamentally different objectives. Taken together, these findings suggest that the gap between real progress and accurately measured progress in RLVR may be substantially larger than the community currently appreciates. We call for more honest accounting practices in RLVR research before the next benchmark milestone is celebrated.
Chat is not available.
Successful Page Load