MemLoRA: Specializing Small Language Models with Expert Adapters for Lightweight Agent Memory
Abstract
Memory is what keeps conversational agents consistent across long, multi-session interactions: at every turn, salient facts are extracted, consolidated, and retrieved as context. These memory operations are repetitive and narrowly scoped, exactly the kind of sub-task Small Language Models (SLMs) are well suited to; yet current systems delegate them to large, often cloud-hosted models, because prompting an SLM directly performs poorly. We introduce MemLoRA, an adapter-based recipe built on Mem0, a widely adopted LLM-based memory system. A single SLM is equipped with lightweight expert adapters, one per memory operation: knowledge extraction, memory update, and memory-augmented generation. Each adapter follows a simple recipe: distill from the teacher by default, and use ground-truth answers when available. MemLoRA outperforms 10× larger baselines on LoCoMo, matches 60× larger models, and runs 15--35× faster, turning memory into a lightweight, self-contained component of an agentic pipeline that no longer depends on a large backbone.