Retrieval-Augmented Memory Is Not Enough: The Monotonic Retrieval Paradox in Long-Term Conversations
Abstract
Conversational validity is non-monotonic: a later utterance can invalidate an earlier commitment. Candidate-local retrieval support moves the other way, since extending the history cannot lower a stored item's score unless its record is rewritten or the scorer reads supersession relations. We call the mismatch the Monotonic Retrieval Paradox (MRP): a superseded commitment loses validity, keeps its support, and can outrank its replacement. Our position: a relevance-only path cannot guarantee current validity unless supersession information is represented and consumed before generation, and where stored state is rewritten the burden moves to verification prevailing evaluations cannot supply. Verification requires a conjunctive four-stage contract, detection through readout, which a paired intervention audit measures stage by stage. An LLM extraction memory escapes score invariance at ingest and meets our supersession marker on all 30 single-event items, yet serves a superseded value in 46\% of eight-relation conversations. A snapshot arm isolates composition failure: all 400 revisions reach the store, and of 37 conversations whose every update is marked, 2 answer correctly. A closed-family, single-slot state machine reaches 100\% at every tested depth. Where the transcript underdetermines validity, a total-variation data processing argument gives an error floor no passive design removes, whatever its context length or store. The position therefore has two levels: an engineering failure where explicit evidence is present, and an observational limit, of unmeasured size on real conversations, where it is absent; no universal open-language impossibility claim is made. We specify the retraction-labelled corpus the surveyed benchmarks do not supply, and release our audit protocol and generators.