Optimizer Memory Schedules for Outscaling the Overtraining Axis
Katie Everett ⋅ Shikai Qiu
Abstract
Optimizer comparisons are often made at a single training horizon, even though deployed language models are increasingly trained far beyond the compute-optimal token budget. We study AdamW, ADANA, Muon, and SOAP across 51M--253M parameter models and overtraining (OT) factors from $1\times$ to $256\times$, independently sweeping the base learning rate at every setting. Muon and SOAP provide large short-horizon gains over AdamW, whereas ADANA improves relative to AdamW as training length increases. Longer horizons also favor longer fixed optimizer memory, but ADANA's scaling advantage persists after tuning AdamW's fixed memory separately at each horizon. With log-time weight decay and a momentum cooldown that shortens memory during terminal learning-rate decay, ADANA's equivalent-OT scaling against AdamW has fitted slopes $1.15$--$1.20$, close to the $2-\kappa=1.15$ effective-clock reference from DANA theory. ADANA begins behind Muon and SOAP but closes these gaps with additional training. These results suggest that training horizon is an essential axis for optimizer evaluation and that scheduled memory provides a distinct route to optimizer outscaling.
Chat is not available.
Successful Page Load