Personalize-then-Store: Benchmarking and Learning Personalized Memory for Long-Horizon Agents
Abstract
Existing large language model (LLM)-based memory systems apply universal, static policies that overlook a fundamental reality: the contexts that are worth storing in memory are different across users. This misalignment wastes limited memory budget on transient interactions while failing to preserve critical context for long-horizon tasks. To address this gap, we investigate an underexplored question: can LLM-based memory systems learn personalized memory policies? We introduce PerMemBench, the first benchmark for evaluating personalized memory systems, featuring multi-year, multi-domain interaction histories across diverse user personas. We further present the first empirical study of memory personalization and propose simple baseline methods. Our empirical study confirms that personalization yields substantial retention gains when the user profile is exactly inferred, yet reveals that accurate profile inference remains an open and critical challenge.