Bridging Compute- and Data-Optimal Pretraining
Tian Qin ⋅ Kimia Hamidieh ⋅ David Alvarez-Melis
Abstract
Classical compute-optimal scaling laws assume an unbounded supply of fresh pretraining data, yet pretraining is increasingly entering the regime where compute is growing faster than high-quality data. We propose *Compute-Data (CD) scaling laws*, a unified framework that bridges *compute-optimal* scaling, in which data scales freely with compute, and *data-optimal* scaling, in which the corpus is fixed and compute can grow unbounded. CD scaling extends classic scaling by introducing a *token effectiveness function* that quantifies how much a *derived* token, produced for instance by multi-epoch repetition or paraphrasing, is worth relative to a fresh one, ranging from a perfect substitute to no value at all. Fitting $\eta$ for two data expansion strategies (multi-epoch repetition, paraphrasing) across model sizes ranging from 14M to 600M parameters on the Dolma-3 corpus, we find that it is far from constant: it depends jointly on model size, the tokens-per-parameter ratio, and the amount of derived data, and saturates as the corpus is expanded. The functional form of $\eta$ implies that substituting compute for data *diminishes* with both model size and data availability, and partitions training into three operational regimes: compute-bound, data-bound, and model-bound, showing that classic compute-optimal allocation is suboptimal across most of the practically relevant regime.
Chat is not available.
Successful Page Load