MoME: Mixture-of-Memory Embeddings for Context-Aware Sparse Lookup
Muchen Li ⋅ Leonid Sigal ⋅ Renjie Liao
Abstract
Scaling large language models efficiently has motivated sparse capacity mechanisms such as Mixture-of-Experts and, more recently, conditional memory: token-indexed embedding tables that augment the backbone with cheap parametric lookups. Existing memory-embedding methods retrieve via a deterministic function of the surface form, which collapses different contextual senses of the same token (\eg, \emph{python} the language vs.\ the animal) into a single fixed entry. We introduce Mixture of Memory Embeddings (MoME), a context-aware memory mechanism that replaces each token's single memory row with a mixture of $M$ slots and uses a learned gate over the hidden state to choose which slots to read at each position. In controlled pretraining experiments across nanochat, Llama-3/MobileLLM, and Qwen3 backbones from $125$M to $0.6$B parameters, MoME consistently outperforms Value Embedding, Engram, and STEM baselines at matched memory \& training budgets, shows a more favorable memory-size scaling trend than Engram on the nanochat backbone, and remains compatible with Engram-style memory under compound scaling. Qualitative routing analyses on polysemous tokens further suggest that the learned mixture exhibits a degree of semantic interpretability, dispatching the same surface token to distinct memory slots under different senses rather than collapsing them into a single fixed entry. All Code and model checkpoint will be open sourced.
Chat is not available.
Successful Page Load