Rethinking Exploration in Modern Mixture-of-Experts
Abstract
Mixture-of-experts (MoE) models have become an important approach to scaling model parameters. Modern MoE architectures increasingly combine two design choices: fine-grained expert segmentation and shared experts. The former increases expressiveness by enabling more expert combinations at a fixed activation ratio, using narrower experts while increasing both the total and activated numbers of experts. We reexamine the interaction between these design choices and find that shared experts can make routing less explorative by reducing the variability of token representations. This conflicts with the greater need for exploration induced by fine-grained expert segmentation, which combinatorially enlarges the routing space. Consequently, the additional expressiveness offered by fine-grained segmentation may not be fully realized in practice. MirrorMoE is proposed to mitigate this tension. It adaptively reduces the exploration space by functionally suppressing unnecessary routing paths on a per-token basis. By reconciling these competing effects, MirrorMoE consistently improves performance across model scales.