Position: We Should Evaluate Agentic Memory as Inference, not as a Tape Recorder
Abstract
Evaluating memory in AI systems is increasingly shaped by how memory is implemented (e.g., retrievers, context windows, parameter updates) and by what the system produces (e.g., accuracy, retrieval recall). Such views risk overlooking a key question for understanding memory: why does a system produce a particular output given its past experience? In this position paper, we posit that memory is fundamentally an inference process in which an agent uses past information to form beliefs and make decisions under uncertainty, and that AI memory problems should be evaluated as problems of inference. We contend that existing fragmented, system-specific evaluations might misdirect research toward narrow fixes rather than transferable and explainable understanding. To make this view operational, we formalize memory-guided behavior as conditional inference over episodic traces, cues, and schemata: we derive an architecture-agnostic taxonomy of memory failure modes comprising schema dominance, trace fragmentation, and provenance collapse, and a suite of intervention-based evaluation criteria that perturb episodic evidence while holding the cue and architecture fixed. We instantiate a retrieval-augmented LLM agent as a proof of concept, and present ablations on the LoCoMo benchmark spanning five memory architectures and six LLM backbones. Our preliminary findings corroborate our view: models which are statistically indistinguishable on task accuracy diverge significantly under intervention, most sharply in provenance. We postulate that this view offers a principled foundation for evaluating and advancing memory competence across AI systems.