Two Levers for Undermined Facts in Agent Memory: What the Store Records and What the Reader Is Told
Abstract
An agent’s memory can retain a fact that is no longer safe to assert. When a value was recorded on account of another fact, and that other fact subsequently changes without a replacement being stated, nothing in the store contradicts the value, yet it can no longer be asserted; we call such a fact undermined. An assistant told in March that a user’s health insurance is provided through their employer, and in August that the user has changed employers, should not report the March coverage as current in September. MEME, a recent multi-session memory benchmark, labels these Absence items and reports that memory systems answer approximately 1% of them correctly. The cause is structural: such systems retire a stored value when a competing value arrives, and no competing value arrives here. We distinguish two interventions that could address this, owned by different parts of a deployment: at write time the store can record why a fact was written, and separately the agent’s system prompt can instruct the answering model, which we term the reader, to act on such records. We show that no published measurement separates them, because MEME serializes the dependency into the stored text it evaluates and its own prompt-optimization ablation held the answering prompt fixed. We therefore re-render MEME’s 130 Absence items at three levels of recorded dependency, from a store that records nothing to MEME’s own serialization, and cross this against whether the reader is instructed, across five reader models. We find the two interventions strongly asymmetric. A three-sentence system prompt naming no entity, question or answer raises the best reader’s accuracy from 0.015 to 0.923, while recording the dependency and leaving the prompt unchanged reaches at most 0.192, or 0.569 where the store also records what follows from that dependency. Neither intervention suffices alone. Because an instruction to express doubt may induce doubt indiscriminately, we reconstruct every item a second time with the change event removed, which leaves the recorded value correct, and score both versions under one rubric that states no expected answer. We measure genuine discrimination for three of the five readers, 0.75 for the best and near zero for the weakest, and find that the reader gaining most on the original task is not among them: it flags 0.55 of values whose grounds never lapsed. We also find the instruction actively harmful where a fact’s grounds have lapsed but its replacement is derivable: erroneous hedges rise from 0.166 to 0.534 and the best reader loses more than half its accuracy. We conclude that the less costly intervention produces the larger gain in accuracy, and that it cannot be deployed unconditionally.