Scaling Arbitrary Architectures and Optimizers with Automatic Parameterization
Shikai Qiu ⋅ Charlie Chen ⋅ Andres Potapczynski ⋅ Martin Marek ⋅ Andrew Wilson
Abstract
Training neural networks at scale requires careful per-layer scaling of initialization, learning rate, and other optimizer hyperparameters, with the right rules depending jointly on the architecture, the optimizer, and which axis (width $D$, depth $L$, batch size $B$, etc.) is being scaled. Current practice is to derive each rule by hand, as in $\mu$P for width, CompleteP for depth, SDE-based and EMA-timescale arguments for batch size, and tailored prescriptions for matrix-preconditioned optimizers; as architectures, optimizers, and scaling axes continue to evolve, this hand derivation is increasingly a bottleneck to combining these advances at scale. We show that these derivations can be systematized by solving a linear system over hyperparameter exponents to satisfy the scaling constraints imposed by each primitive instruction of the training program, analogous to dimensional analysis in physics that enforces matched units on both sides of an equation. We develop \emph{automatic parameterization} (AutoP), a system that reads these rules off a traced training program and solves the resulting system to assign a scaling exponent to every hyperparameter, and provide a concrete JAX implementation. From a few primitive rules, AutoP recovers $\mu$P for width on Transformers, Monarch-structured Transformers, mixture-of-experts, and ResNets under Adam, SignSGD, Muon, and AdaMuon; CompleteP for depth; and the known batch-size prescriptions for the learning rate and AdamW weight decay. It also identifies new rules for the recurrence count in looped transformers, the context length in linear-attention and MLP-Mixer models, and the batch-size scaling for Muon's learning rate, all of which we demonstrate as empirically beneficial. Analogous to automatic differentiation, our results suggest that robust hyperparameter scaling rules can be automated, freeing practitioners to scale up new architectures and optimizers without rederiving the underlying theory or risking subtle errors that quietly cost compute at scale.
Chat is not available.
Successful Page Load