Towards Principled Fine-Grained MoE Expert Pruning via Pseudo-Boolean Approximation
Abstract
Mixture-of-Experts (MoE) language models improve parameter efficiency by activating only a small subset of experts per token, but all experts must still be stored at inference time, making memory a key deployment bottleneck. Expert pruning is therefore a practical post-training compression strategy, but existing methods face a fundamental tension: search-based approaches attempt to capture expert interaction effects through joint optimization, yet become intractable for modern fine-grained MoE models; score-based approaches, by contrast, ignore such effects, but remain efficient and often surprisingly effective. To understand this tension, we formulate expert pruning as a constrained pseudo-Boolean optimization problem and empirically analyze the interaction effects induced by pruning. Our results show that cross-layer interaction effects are relatively more important, since pruning earlier layers changes the hidden-state distribution for later ones, whereas within-layer effects are often well approximated by additive singleton damages. Based on this observation, we propose \textbf{MoE-PBA}, a prefix-conditioned pruning method that prunes layers sequentially under the hidden states produced by the already-pruned prefix. Across four representative fine-grained MoE language models with distinct architectures and parameter scales ranging from 7B to 30B, MoE-PBA achieves a substantially better accuracy--efficiency trade-off than search-based methods and outperforms the strong score-based baseline on diverse downstream benchmarks.