ALIBI: Auditing whether memory-augmented agents retrieve for the reasons they claim
Abstract
Benchmarks for long-term memory in AI systems typically score the final answer rather than the retrieval that produced it, or the rationale a system gives for retrieving one memory over another. A system can therefore succeed by flooding its context with loosely related memories, or by selecting the right memory and explaining that choice with a story unrelated to what actually drove it. We introduce ALIBI, a protocol that evaluates the justification of memory selection rather than only its outcome. Cases are constructed so that the reason one memory should beat its distractors is a known structural label, which turns justification scoring into classification without an LLM judge, and faithfulness is measured by minimal counterfactual edits to the memory store: a system that says it chose an entry for its recency should change its choice when the dates are swapped. Placebo edits that leave the winning memory unchanged control for generic prompt sensitivity. We instantiate the protocol on 106 human-verified cases across four sources: a deterministic synthetic control, the diaries of Samuel Pepys (1660-1669) and W.,N.,P. Barbellion (1903-1917), and multi-session dialogues. Four systems are evaluated under a single open-weight backbone: two retrieval strategies, a contamination baseline, and one deployed memory system. The retrieval strategies cite the correct memory in about three quarters of cases but name the correct reason in barely a quarter; they respond to the deletion of a cited memory (74-75\%) but almost never to the inversion of dates (8-16\%), while placebo sensitivity stays at 1-3\%. The gap widens with corpus realism: justification accuracy falls from 69\% on synthetic cases to 2\% on Pepys, while selection stays high, and the pattern holds for the deployed system as well. Our data, code and annotation protocol are included as supplementary material and will be released publicly.