Target-ESS Tempered Policy Optimization (TETPO): Smooth Variance Control via Bridge Importance Sampling
Shayan Mohajer Hamidi ⋅ Kayhan Behdin ⋅ Wenhui Zhu ⋅ Aida Rahmattalabi ⋅ Jincheng Cao ⋅ Jelena Markovic-Voronov ⋅ Alborz Geramifard
Abstract
Asynchronous reinforcement-learning post-training of large language models is computationally efficient but inherently off-policy: the learner updates on rollouts generated by a stale policy, while correcting this mismatch requires importance weights that can be extremely heavy-tailed. A small number of trajectories may then dominate an update, collapsing the effective sample size (ESS) and increasing the variance and instability of the resulting gradient estimates. Existing stabilizers apply fixed interventions regardless of the realized severity of weight concentration, either discarding selected samples or reweighting them without adapting the correction strength. We introduce \emph{Target--ESS Tempered Policy Optimization} (TETPO), which replaces each importance weight $\rho$ by $\rho^\alpha$ for a data-adaptive exponent $\alpha\in[0,1]$: $\alpha=1$ preserves the original importance weights, whereas $\alpha=0$ assigns uniform coefficients. For each prompt group, TETPO selects the largest $\alpha$ whose tempered weights satisfy a target ESS fraction $\kappa$, applying little tempering to healthy groups and stronger tempering when weight degeneracy is severe. At the population level, the tempered coefficients are exact importance weights for a geometric bridge between the sampler and learner policies. The exponent is efficiently found by one-dimensional bisection, while the resulting weights satisfy deterministic bounds on normalized-weight dispersion and weighted-gradient norm that are independent of policy drift, together with a fixed-$\alpha$ bound on bridge-target discrepancy. Power tempering also introduces no explicit zero-weight mask. On mathematical-reasoning benchmarks under high staleness, TETPO outperforms strong on- and off-policy baselines. The code is available at \url{https://anonymous.4open.science/r/tetpo-offpolicy-rl-1A41}.
Chat is not available.
Successful Page Load