NGN-Scion: Making Norm-Constrained LMO Optimizers Robust to the Learning Rate
Abstract
Pre-training modern deep learning models demands substantial computational resources, so an improperly chosen peak learning rate can be as costly as a failed run. Because the optimal learning rate varies with model size, batch size, and task, identifying it through large-scale sweeps is prohibitively expensive, which has motivated adaptive algorithms that set the learning rate from statistics such as gradient norms and training losses. Prior work has focused primarily on Adam-type methods, leaving geometry-aware optimizers such as Muon and Scion comparatively underexplored. We adapt the NGN step-size rule to the Scion optimizer, improving its robustness to the choice of Frank--Wolfe step size, and call the resulting algorithm NGN-Scion. We analyze the stability of its step size relative to several baselines and evaluate it empirically. In Chinchilla-optimal pre-training with models from 70M to 410M parameters and in Qwen2.5 fine-tuning from 0.5B to 3B parameters, NGN-Scion matches the best-tuned baseline at every scale while being substantially less sensitive to the learning rate. Code is available at https://anonymous.4open.science/r/scion-ngn/README.md.