Reformulate LLM Reinforcement Learning for Stable Training under Black-box Discrepancy
Abstract
Reinforcement Learning (RL) has emerged as a pivotal post-training paradigm for Large Language Models (LLMs), yet it frequently suffers from unpredictable training collapses. Recent findings attribute these failures to a hidden train-inference discrepancy (or mismatch), such a discrepancy will increase as training goes on, stemming from the disparate underlying engines and precisions required to balance generation throughput and training fidelity. Such Existing mitigations either sacrifice numerical stability or rely on heuristic masking that fails to explicitly optimize the ultimately deployed policy. In this paper, we discover that training policies inherently possess the capability to heal this common and harmful discrepancy. To operationalize this, we transition the standard RL objective into a Discrepancy-Constrained Markov Decision Process (\texttt{DCMDP}). In order to practice this new paradigm at the algorithmic level, we first introduce a robust, trajectory-level geometric penalty that provides black-box feedback of inter-policy deviations, enabling autonomous self-correction under any specific trigger, e.g., infrastructure level or model architecture level. Furthermore, capitalizing on our empirical discovery of a \textit{discrepancy tolerance region} in various models, we employ an adaptive performance-discrepancy balancing mechanism that penalizes the policy only when deviations exceed a safe boundary, achieving stable dual-objective optimization. Our approach not only eradicates mismatch-induced collapses and shatters performance bottlenecks, but crucially unlocks a heterogeneous training paradigm—leveraging high-fidelity training environments to natively optimize LLMs for low-cost, resource-constrained deployments.