Why Routers Freeze: Infinite Width Learning Dynamics for Mixture of Experts
Anish Dhir ⋅ Volkan Cevher ⋅ Leena Chennuru Vankadara
Abstract
Mixture-of-Experts (MoE) architectures are central to large-scale deep learning, relying on sparse execution and expert specialisation. Despite their success, their large-scale training dynamics remain poorly understood, and training is often unstable. While scaling limits such as infinite-width theory have clarified the behavior of dense models, the scaling behavior of MoEs remains largely unexplored. To this end, we derive the infinite-width limits of MoE architectures under fixed number of experts with both soft and Top-$K$ routing, trained with SGD and Adam, using Tensor Programs. Our results show that under the Standard Parameterisation (SP), router updates freeze after one step of training in both soft and Top-$K$ routing MoEs. This provides, at least in part, a principled explanation for the widespread use of auxiliary losses such as load-balancing or $z$-losses in MoE training. In dense networks, Maximal Update heuristics allow for stable and non-vanishing feature updates. However, in MoEs the Maximal Update heuristic with soft-routing and softmax gating causes the experts to collapse to the same distribution and router gradients to vanish. While sigmoid allows non-trivial router evolution, they still fail to prevent expert collapse. Thus fixed expert soft-routing MoEs do not admit a parameterisation that simultaneously ensures stability, feature learning, and expert specialisation in the infinite-width limit. We then show that, beyond computational efficiency, Top-\(K\) also acts as a symmetry breaking mechanism allowing for both feature learning and specialisation. We empirically validate these predictions, observing close agreement between finite-width training dynamics and the predicted scaling behaviour.
Chat is not available.
Successful Page Load