EndlessMemory: Understanding Peak Capacity in Passive Long-Term Memory Systems for LLM Agents
Abstract
Long-term memory is a key component of test-time continual learning agents, enabling agents to accumulate and reuse experiences over extended interactions. We ask whether this benefit can be sustained as agents accumulate increasingly large amounts of memory. To study this question, we introduce EndlessMemory, a framework for measuring how retrieval and downstream answer quality change as an agent's stored experience grows. We study both controlled synthetic expansion and naturally ordered multi-session histories. Using three representative passive memory architectures, including BM25, dense embedding retrieval, and RAPTOR, we identify a system-dependent peak capacity phenomenon under controlled memory scaling: within these configurations, increasing stored experiences initially improves agent performance but eventually causes degradation beyond a system-specific scale. Through oracle--distractor controls, we show that contextual interference and evidence position can reduce performance even when answer evidence is available. Our findings reveal a degradation pattern associated with passive memory accumulation rather than a universal upper bound on agent memory capacity. We also explore whether active memory management strategies, such as selective retention and consolidation, can mitigate this degradation. Finally, we provide information-theoretic and geometric analyses as hypotheses to interpret architecture-specific degradation patterns, while distinguishing memory scaling effects from the LLM's inherent sensitivity to long, distractor-heavy contexts.