Visual Expert Skipping for MoE-MLLMs via K-Armed Bandit based Expert Estimation
Abstract
Despite great success, existing MoE-MLLMs still suffer from substantial visual expert redundancy, making visual expert skipping a critical solution. However, deriving an effective and fixed skipping policy is still an open problem, mainly due to the exponentially large layer-wise policy space and the dependency between expert allocations across layers. In this paper, we propose a novel and training-free approach for MoE-MLLMs, termed \emph{K-armed bandit based expert redundancy estimation} (\textbf{KAB-MoE}). KAB-MoE defines the retained visual expert number at each MoE layer as a discrete action, and then conducts efficient one-shot RL estimations to obtain the layer-action reward matrix. Based on this reward matrix, KAB-MoE can quickly derive the optimal skipping policy that satisfies the predefined computation budgets, \emph{e.g.}, expert skipping ratio. The obtained skipping policy can be directly applied to MoE-MLLMs for all examples, which not only facilitates high-throughput deployment but also contributes to practical acceleration. Extensive experiments on Kimi-VL-A3B and Qwen3-VL-MoE show that KAB-MoE effectively accelerates MoE-MLLMs while keeping strong multimodal performance, \emph{e.g.}, retaining 97.09\% average performance on Kimi-VL-A3B under the skipping ratio of 83\%. Moreover, KAB-MoE also achieves competitive or superior performance over dynamic expert skipping SOTAs with practical inference speed-up. \textbf{Our code} is provided in the supplementary materials.