Neural Scaling Laws in Particle Jets
Abstract
Large language models (LLMs) have demonstrated that scaling, driven by compute, can yield dramatic improvements in performance, with current state-of-the-art systems reaching up to one trillion parameters. In contrast, state-of-the-art jet flavor classifiers in particle physics remain in the tens of millions of parameters. Scaling laws offer a framework for predicting how much particle physics models benefit from systematic increases in capacity and training data. In this work we study the behavior of increasingly larger jet classifiers, deriving scaling laws under optimal use of compute budget and limited datasets. We also identify the regimes where “double descent” emerges. The emerging scaling behavior points to a clear need for significantly larger datasets and models to saturate available compute and approach the limit of achievable performance. We therefore train a classifier on a new dataset 20 times larger than the current state-of-the-art (7.7 billion jets) and find that the observed performance follows the scaling laws prediction. This motivates continued integrated efforts in data generation, model scaling, and inference infrastructure for large-scale models in the physical sciences.