When Load-Balancing Goes Too Far: Expert Pruning in Over-Dispersed Mixture-of-Experts Models
Berkcan Kapusuzoglu ⋅ Connor Pryor ⋅ Sangwoo Cho ⋅ Supriyo Chakraborty ⋅ Shi-Xiong Zhang ⋅ Sambit Sahu ⋅ Milind Naphade
Abstract
Mixture-of-Experts (MoE) models are expensive to deploy because the entire expert population must sit in memory even though each token activates only a few of them, so total parameter count, not per-token compute, sets the GPU and serving budget. Expert pruning attacks this footprint by removing low-importance experts flagged by the router, which assumes router probabilities are a reliable importance signal. We show the assumption breaks under $\textit{over-dispersed routing}$, a regime induced by aggressive load-balancing during training (e.g., $\lambda_{\text{aux}}{=}0.9$ in gpt-oss-20b), where tokens spread almost uniformly across experts and importance signals collapse. Perplexity stops tracking downstream accuracy, and pruning forces a capability trade-off: on gpt-oss-20b, activation-aware scoring keeps mathematical reasoning intact but sacrifices knowledge-intensive science by 18 GPQA points. Minimax Expert Score Allocation (MESA) scores experts to protect whichever domain a candidate pruning plan currently hurts most. At 25\% expert pruning it gives the smallest worst-case degradation across domains, and it generalizes across scale and architecture: within the gpt-oss family to 120B, to a different architecture (Gemma-4-26B-A4B, $+14.3$pp GPQA over the strongest baseline), and to a third, base MoE family (OLMoE-1B-7B, retaining 92.8\% of unpruned commonsense accuracy). Under standard routing (Mixtral-8x7B-Instruct) the trade-off disappears, so what MESA fixes is a property of the over-dispersed regime rather than of pruning in general.
Chat is not available.
Successful Page Load