Useful Memories Become Faulty When Continuously Updated by LLMs
Abstract
Learning from past experience requires forming abstractions that can be reused in future problems. Recent work on agentic-memory systems has explored a practical route to continuously improving LLM agents after deployment without parameter updates: a textual abstraction is derived from each history and stored in memory, then continuously updated with more interactions. Yet we show that such textual abstractions produced by today's LLMs are often faulty, even when derived from useful experiences. Over time, the memory becomes less useful and even harmful. More surprisingly, even when abstracting from ground-truth solutions, GPT-5.4 fails on 54% of a set of ARC-AGI problems it had previously solved without memory. We attribute this to the iterative update process: the same trajectories yield qualitatively different memories under different update schedules, while an episodic-only control over those trajectories remains competitive with the consolidators we test---pinning the failure on the consolidation step, not the underlying experience. In our ARC-AGI Stream environment designed to trace memory consolidation behavior, when agents are given additional actions, they tend to keep episodic memory and double the accuracy of their forced-consolidation counterparts; removing consolidation entirely (episodic management only) matches this auto mode. These results suggest that current LLMs should not recursively rewrite their own experience into stable long-term knowledge. Robust agent memory should treat raw episodes as first-class evidence and make abstraction selective, delayed, and explicitly gated rather than mandatory after every interaction.