Scaling Laws for Multimodal Data Mixtures
Abstract
Frontier AI systems are increasingly natively multimodal, jointly pretrained on text, vision, and audio. Yet the data mixtures behind them, how much of each modality to include, and at what cost to the others, are largely undocumented. Public scaling recipes cover text but do not tell practitioners how modalities interact during joint training. We derive data-mixing scaling laws for text, vision, and audio, using Mixture-of-Experts (MoE) models, which are the emerging default for multimodal pretraining. We extend the Chinchilla law to multimodal MoE pretraining. Our law predicts a target modality’s validation loss under three effects. Directional cross-modal transfer captures when a source modality helps or hurts a target. Mixture-dependent capacity sharing captures how modalities compete for finite MoE capacity. Repetition saturation captures the diminishing returns of re-passing scarce non-text data. Our law offers practitioners a principled tool to plan mixtures and estimate the cost of adding a modality, without exhaustive trial-and-error.