Primitives and Paths: Predicting What Post-Training Can Do from What the Base Model Already Has
Anqi Qu
Abstract
Does reinforcement learning teach language models new capabilities, or merely amplify ones acquired during pre-training? The literature offers evidence for both: pass@$k$ studies suggest RL reweights latent behaviours, while some controlled composition studies demonstrate new compositional or exploratory gains. We argue that this question has no answer independent of the capability state of the base model, and propose a two-dimensional state space for post-trainability. The first axis, effective full-solution support, measures whether complete successful trajectories are recoverable under the base policy at feasible sampling budgets. The second axis, atomic primitive coverage, measures whether the constituent skills of a solution are individually available when the burden of composition is removed. These axes distinguish three qualitatively different post-training regimes -- amplification, composition-limited, and acquisition-limited -- each predicting a distinct empirical signature and calling for a distinct intervention. There also exists a fourth diagnostic region indicating measurement failure rather than a learning mechanism. We propose a measurement protocol for locating a model–task pair in this space before post-training, including an internal-consistency check derived from the fact that primitive coverage mechanically implies a floor on solution probability. The framework reframes mid-training as preparation for post-training: its role is to move models into states from which RL can operate effectively.
Chat is not available.
Successful Page Load