Adaptive Decoding Methods for Token-Efficient Reinforcement Learning
Zhaonan Li ⋅ Zelei Cheng ⋅ Amritansh Mishra ⋅ Michael Lee ⋅ Mann Patel ⋅ Sambit Sahu ⋅ Ben Zhou ⋅ William Campbell
Abstract
Rollout decoding affects both which trajectories are used for reinforcement learning with verifiable rewards (RLVR) and how many completion tokens are spent collecting them. We study whether candidate-token support should remain fixed or widen during training in the context of rollout sampling efficiency. Our main schedule moves from top-$K=10$ to $K=20$ and then to full support as the attempted-token budget increases. In a 4B code-RLVR setting, the schedule has a higher interpolated 150M-token validation value than its matched fixed-$K=10$ continuation in all three runs, with a mean gain of $1.80{\pm}0.42$ points. The schedule also performs comparably to a separately trained fixed-$K=20$ baseline, the strongest fixed-support setting in our sweep. Follow-up experiments suggest that late widening to full support is the most promising variant, although the evidence remains noisy, and mixed-reward-group diagnostics do not clearly explain the gain. An exploratory single-run test replay yields identical point estimates for the schedule and fixed $K=10$. Overall, our results show that rollout support is an important control for sampling efficiency in RLVR and motivate further study of adaptive support schedules.
Chat is not available.
Successful Page Load