μLM: Rethinking Sub-100M Language Models through Memory-First Design
Zijie Chen ⋅ Guiyun Fan ⋅ Haiming Jin
Abstract
Ubiquitous IoT devices, such as companion toys, wearables, and micro-robots, are increasingly embracing language models, yet current cloud-based deployments suffer from latency, privacy, and weak-network reliability issues. These resource-constrained devices typically cap deployable model size below 100M parameters, while existing sub-100M models fall short in basic language and commonsense capabilities. In this extreme regime, we argue that the memory bottleneck supersedes the traditional compute bottleneck as the binding constraint, calling for memory-first design that maximizes model capability under a fixed parameter budget. Notably, the stringent budget binds both what the model stores and what it can afford to learn, demanding architecture-data co-design. In this paper, we propose $\mu$LM, a memory-first sub-100M language model framework instantiating this co-design. For architecture, we present DMLA, which preserves latent attention's compact cache while relieving its feature-squeezing bottleneck under aggressive compression via explicit content-position latent decoupling. For data, we present $\mu$Mix, a capacity-aware multi-stage curriculum tailored to sub-100M models, prioritizing high-quality natural language and commonsense data aligned with embedded scenarios over brute-force token scaling. $\mu$LM-70M surpasses all sub-300M baselines and rivals several sub-1B models on standard benchmarks; even at 9M, it remains competitive. Fine-tuned into a tool-calling agent, $\mu$LM outperforms same-scale baselines and achieves a 97.2\% end-to-end executable rate, presenting a promising direction toward practical sub-100M language modeling.
Chat is not available.
Successful Page Load