Modality-Aware Expert Pruning for MoE-Based Multimodal Large Language Models
Abstract
Mixture-of-Experts (MoE) enables multimodal large language models (MLLMs) to scale capacity efficiently, but expert parameters dominate memory—accounting for over 90% of model size in modern MoE-MLLMs (e.g., 54 GB out of ∼58 GB in BF16 for InternVL3.5-30B-A3B and Qwen3-VL-30B-A3B)—even though only a few experts are activated per token. Expert pruning has emerged as an effective compression strategy, but existing methods aggregate importance uniformly across tokens, ignoring that the same expert may serve different roles for vision versus language processing. To address this, we propose MAEP (Modality-Aware Expert Pruning), a training-free framework that decomposes expert importance by modality in a single forward pass. MAEP introduces (1) Delta-H Output (∆H), a redistribution-aware metric measuring the MoE output change when an expert is removed, and (2) low-attention vision tokens as a vision-side pruning signal that is dataset-stable: its expert-importance ranking remains stable across different datasets. Text-side importance serves as a preservation signal while vision-side importance serves as a pruning signal, and experts are ranked globally across layers. Experiments on three MoE-based MLLMs across multimodal and text benchmarks show that MAEP outperforms modality-agnostic baselines on the majority of multimodal benchmarks while remaining competitive on text-only tasks, reducing the BF16 expert weight footprint by ∼27 GB (∼47% of total memory) so that the expert-weight footprint fits within a single 40 GB GPU.