When Are Compositional Problems Learnable from Verifiable Rewards?
Abstract
Many problems are inherently compositional: solving them requires a sequence of intermediate decisions that jointly determine the final answer. A central question is when such a compositional structure can be learned from outcome-level feedback alone, where supervision of the intermediate steps is unavailable. We study this question for autoregressive models trained using reinforcement learning with verifiable rewards (RLVR). We identify the \emph{task-advantage ratio}, a joint property of the task and the model, which measures whether intermediate decisions present an advantage in reaching a correct final solution. We show that this ratio governs learnability in our setting: when the advantage is present, RLVR efficiently learns the target composition, while when it is absent, training can converge to suboptimal compositions. We further show that the required advantage arises naturally in several structured problems, but may depend critically on the quality of the initial model. Our results help to clarify when compositional problems can be learned from final rewards alone.