Why Did It Forget? Diagnosing Retrieval Failures in Long-Term Conversational Memory Agents
P Sivadhanushya
Abstract
Long-context agents increasingly rely on persistent conversational memory, but reliable memory requires more than storing information: agents must discover, retrieve, and correctly use the right evidence. Yet long-term memory evaluations often emphasize final-answer correctness, obscuring where the memory pipeline actually failed. We introduce the Retrieval Failure Taxonomy (RFT), a diagnostic framework that separates five core failure modes and two refinements spanning pre-retrieval, retrieval-strategy, and post-retrieval generation failures. A blinded two-annotator study ($N=35$) yields 77.1% exact agreement (Cohen's $\kappa=0.734$), while adjudication identifies two boundary refinements and one case outside the current taxonomy. We then evaluate all 500 LongMemEval instances under three retrieval configurations and two answering-model families, yielding 3,000 primary evaluations. Expanding accessible summary context from 35k to 75k characters significantly improves retrieval recall, F1, and final-answer correctness in both models. By contrast, the keyword-search configuration significantly improves retrieval recall and F1 without significantly improving final-answer correctness. This retrieval-correctness dissociation replicates across both model families, showing that better evidence access does not necessarily yield better answers. More broadly, RFT provides a way to evaluate long-context agents not only by whether they fail, but by where in the memory pipeline the failure originates.
Chat is not available.
Successful Page Load