Permissive Data Is Not the Bottleneck: Front-Loading Reasoning Data Enables Strong Post-Training
Ali Elganzory ⋅ Harsh Raj ⋅ Marianna Nezhurina ⋅ Victor May ⋅ Van Khue Nguyen ⋅ David Salinas ⋅ Huu Nguyen ⋅ Jenia Jitsev
Abstract
Permissively licensed pre-training corpora make open foundation model research lawful, auditable, and shareable, but they are widely assumed to cost capability. Recent work reduced the gap in pre-training by front-loading instruction and reasoning data. This leaves post-training unclear and two factors entangled: the licensing of the source pool and the composition of the mixture. We separate them with a $3 \times 2$ factorial design at 1.7B parameters and 300B pre-training tokens, crossing three web dominant sources (a permissive pool, Nemotron-CC-HQ, and FineWeb-Edu) with the presence or absence of the same instruction-and-reasoning subset in the pre-training. Every resulting model undergoes an identical long-context extension and Tulu3 post-training pipeline under same evaluation protocol with three decoding seeds. We show that front-loading drives strong improvement, raising the evaluation suite average in all cells by $10.3$ to $19.2$ points, with gains of $+23.0$ points on code, $+15.6$ on mathematics, and $+10.4$ on instruction following. Licensing does not make a difference in the front-loading setting. Once both pools are front-loaded, the permissive dataset is statistically indistinguishable from both non-permissive datasets at 4K context after post-training and is ahead by $4.7$ to $5.0$ points at 16K, with front-loading dominating the boost irrespective of licensing type. Receptivity to a later reasoning stage is likewise a property of the front-loading reasoning. Removing OpenThoughts3 from a 300B pre-training run and then applying identical reasoning post-training drops the suite average from $33.6$ to $23.9$ and returns AIME24 and AIME25 to zero. The benefit is not specific to post-training on the same corpus that was front-loaded, and appears also when post-training with other reasoning corpora not present in pre-training.
Chat is not available.
Successful Page Load