Towards Optimal Pre-training Data Mixtures for Empowering Reinforcement Learning in LLMs
Xue Han ⋅ Qian Hu ⋅ Yitong Wang ⋅ WenChun Gao ⋅ Lianlian Zhang ⋅ Qing Wang ⋅ Lijun Mei ⋅ Zhaoxuefeng ⋅ Junlan Feng
Abstract
Optimal pre-training data mixtures are vital for Large Language Models (LLMs), yet post-RL reasoning performance varies significantly across base models. This suggests that identifying "RL-friendly" pre-training data mixtures is a critical but under-explored prerequisite. This motivates us to ask: Which specific data mixtures in pre-training enhance post-RL effectiveness, and how can we evaluate a base model's suitability for RL fine-tuning? Through large-scale experiments with proxy models, we observe that (1) domain-specific data correlation strongly predicts post-RL performance, and (2) the average perplexity of the base model on $k$ RL-generated responses, $\text{avg}(k\text{-ppl})$, is a more stable and correlated predictor than traditional metrics. Building on these insights, we propose BridgeRL, a framework that treats finding optimal mixing ratios as a regression task, using $\text{avg}(k\text{-ppl})$ as the fitting objective. We train 1M/60M-parameter proxy models for regression and scale the findings to 1B/3B-parameter models. Results across online and offline RL algorithms demonstrate that BridgeRL-optimized mixtures consistently yield superior base models for RL fine-tuning.
Chat is not available.
Successful Page Load