Post-training Trade-offs Under Limited Compute Budgets
Abstract
Post-training for Large Language Models commonly combines supervised fine-tuning (SFT) with reinforcement learning from verifiable rewards (RLVR), but how compute should be allocated between these stages remains unclear. We formulate post-training as a compute allocation problem and ask how the optimal SFT-RLVR split changes with the available budget. We post-train Llama 3 and Gemma 3 models on mathematical reasoning tasks under low, medium, and high compute budgets, varying the share allocated to SFT and RLVR. We find that low compute budgets favor allocating most or all compute to SFT, while higher budgets favor dedicating an increasing share to RLVR. Similar trends are also found for instruction following tasks. Nevertheless, RLVR-only training remains less effective than combining SFT and RLVR under a fixed compute budget. This contrasts with prior findings that SFT can limit peak performance after RLVR and highlights a distinction between maximizing the RLVR performance ceiling and maximizing performance under fixed compute.