Spectrum-Aware Expert Merging in Latent Space
Abstract
Mixture-of-Experts (MoE) models scale model capacity efficiently by activating only a small subset of experts for each input, while still requiring all parameters to be stored in memory, resulting in substantial memory overhead. Expert merging technique has recently emerged as a promising compression paradigm that reduces this overhead by consolidating multiple experts into a smaller set. However, existing approaches have two key limitations: (i) direct convex interpolation in parameter space is not generally guaranteed to yield meaningful interpolation of expert functions; and (ii) estimating expert similarity or merging coefficients often relies on calibration data, introducing dependence on the calibration distribution. To address these limitations, we propose Spectrum-Aware Latent Merging (SALM), a calibration-free expert merging framework that occurs in a structured latent representation of expert weights using a VAE. SALM performs geometry-aware interpolation in the learned latent space and decodes the resulting representation into a merged expert, while leveraging latent spectral structure to determine expert grouping and merging coefficients. This enables principled, calibration-free expert merging without additional data. Across all evaluated settings, SALM achieves the highest average performance on generation tasks.