Temporal Mixture-of-Experts: Serving Memory that Scales with Active Parameters
Noah Cylich ⋅ Mohsen Fayyaz ⋅ Henry Ndubuaku ⋅ Nanyun Peng
Abstract
Sparse Mixture-of-Experts (MoE) models activate only a small fraction of their parameters per token, yet serving one still requires the whole expert pool in fast memory, making inference impractical on memory-constrained local devices. We introduce Temporal MoE, which makes expert locality a training-time constraint. Our rolling-residency router keeps only the $k$ active experts of each layer resident in RAM and swaps at most one expert per layer per token, so the incoming expert streams in behind the compute of the re-used, resident ones and serving memory scales with active rather than total parameters. We pretrain models across isoFLOP sweeps from $10^{16}$ to $10^{19}$ FLOPs, and the temporal models recover 72–82% of the MoE-over-dense performance gains for both cross-entropy and downstream accuracy at compute-optimal sizes while being 5$\times$ more memory efficient. Router probes demonstrate that the temporal router focuses more on context rather than token identity and spreads each expert's use more evenly across its token stream. Our llama.cpp implementation of Temporal MoE serves an 11B-scale model using 5.1$\times$ less memory while retaining 0.70–0.83$\times$ all-resident decode speed over 1k–4k context lengths, both on an RTX A6000 and on a Pixel 10a, where the model otherwise does not fit in memory. We also apply this constraint to released instruct MoE models and find that when the same fraction of experts are kept resident, sparser models are more robust to the constraint. Two GPU-hours of distillation remove most of the remaining penalty for gemma4-26B and Qwen3.5-35B. Previously proposed serving caches preserve quality with half the experts resident but degrade at our smaller residency budgets. These results show that Temporal MoE can move the quality–memory frontier for local inference by training expert locality into the model.
Chat is not available.
Successful Page Load