TIC-GRPO: Provable and Efficient Optimization for Reinforcement Learning from Human Feedback
Lei pang ⋅ Jun Luo ⋅ ruinan Jin
Abstract
Group Relative Policy Optimization (GRPO) is a critic-free reinforcement learning algorithm for fine-tuning large language models, but its token-level importance sampling mechanism creates a subtle variance bottleneck. We separate this bottleneck into two mechanisms. First, the standard GRPO clipping rule is incomplete: on the negative-advantage branch, it can pass uncontrolled upper-tail importance weights. This motivates up-only clipping as a standalone correction. Second, up-only clipping alone is not enough if the update remains token-level: the clipped token-weighted score increments need not be martingale differences, so cross terms survive in the squared update. This motivates trajectory-level importance correction. We propose Trajectory-level Importance-Corrected GRPO (TIC-GRPO), which follows this progression: it first caps upper-tail ratios by up-only clipping and then replaces token-level importance ratios with a single trajectory-level probability ratio. The latter restores the measure-change identity needed for trajectory-level martingale cancellation. Our variance theory gives upper bounds, hard-instance lower bounds, and induced convergence lower bounds that separate original GRPO, the token-level up-clipped comparator GRPO$_2$, and TIC-GRPO. Experiments on math reasoning and coding benchmarks further confirm that TIC-GRPO improves optimization stability and final performance.
Chat is not available.
Successful Page Load