Strong Post-Training from Permissive, Reasoning-Dominant, Web-Scale Pretraining
Abstract
Permissive, auditable pretraining corpora enable lawful, responsible and transparent open foundation model research and development, facilitating standardization and common progress. Recent work such as MixtureVitae has shown that permissive corpora can also yield strongly competitive base models, but no controlled comparison has measured whether they still remain competitive after post-training. We post-train 1.7B base models pretrained on 300B of MixtureVitae permissive tokens and compare them against strong compute and token matched non-permissive reference baselines (Nemotron-CC-HQ, FineWeb-Edu and DCLM) and against open weights baselines trained at substantially larger pretraining compute (SmolLM2 1.7B, Qwen2.5 1.5B, Qwen3 1.7B). All models are post-trained with the same pipeline: Tulu3 supervised fine-tuning and direct preference optimization, followed by OpenThoughts3 reasoning training. Post-trained MixtureVitae models match or outperform the matched non-permissive reference on reasoning and instruction-following benchmarks, and maintain solid general language understanding performance. They also remain competitive with strong open weights baselines, which use over an order of magnitude more pretraining compute. Ablations confirm that the substantial reasoning and instruction subset of MixtureVitae drives this result: removing it drastically weakens the post-trained model under the same recipe. Together, these findings show that permissive pretraining does not preclude strong post-training, strengthening the case for legally safe and reproducible open foundation model research and development at frontier performance levels. Data, models, and code to reproduce the experiments will be open-sourced.