What is the role of Learning Rate in a Mixture-of-Experts ResNet
Mateusz Pyla ⋅ Igor T Podolak ⋅ Stanislaw Jastrzebski
Abstract
Learning rate is the most crucial parameter guiding the dynamics of neural network training. In Mixture-of-Expert models, the role is extended to control which expert each input is sent to. We train roughly $1900$ MoE ResNets on CIFAR-10 and CIFAR-100 with SGD, AdamW and Muon, and measure what the learning rate $\eta$ and the load-balancing coefficient $\alpha$ each do to that decision. Without a balancing loss the router is collapsed from initialisation and stays collapsed: the busiest expert takes every input at every step size we measure, under various optimisers, while test error varies by up to $0.08$. Introducing a LayerNorm on the router input removes total collapse but does not totally balance the router; decreasing the need of $\alpha$ by about a factor of ten. On the other hand, the balancing costs little accuracy, but it leaves a router whose decisions are easy to move. Under CIFAR-10-C a balanced router changes $60\%$ of its routing decisions at severity $5$ against $29\%$ for an unbalanced one, and a worst-case input perturbation too small to see flips $12$--$15\%$ of them against $0.1\%$ for random noise of the same size. The learning-rate ordering under corruption reverses between the two regimes.
Chat is not available.
Successful Page Load