Route-Consistent Adaptation for Stable Quantization of Mixture-of-Experts Models with Theoretical Guarantees
Abstract
Quantizing Mixture-of-Experts (MoE) models is difficult for two coupled reasons: expert quantization error perturbs expert outputs, and sparse routing amplifies small perturbations into discrete expert selection changes. This coupling produces progressive output drift that is not captured by dense model quantization analyses. To address this gap, we propose a unified quantization framework that jointly handles both sources of error. First, we formulate compression-induced routing inconsistency as a form of training-inference router mismatch and use this connection to motivate limited routing replay, which replays BF16 routes during quantization-aware fine-tuning (QAT), freezes router parameters, and updates only expert-side adapter parameters. We then provide a standard convergence guarantee for the resulting fixed-route QAT objective, together with a routing-consistency bound under explicit router-margin conditions. Second, for expert restoration under a fixed rank budget, we propose load-aware low-rank compensation that assigns larger compensation rank to heavily loaded experts. Under this allocation rule, we prove monotone strict decrease of weighted compression error and characterize optimal rank assignment by marginal load-weighted spectral gains. Finally, experiments on the MXFP4 quantization of Mixtral, DeepSeek, and Qwen3 MoE models show that our method improves perplexity and downstream accuracy over prior methods. We also report system efficiency on Qwen3-30B, where uniform W4A8 serving accelerates decoding and our adapter path preserves most of that acceleration.