When Completion Goes Wrong: Source Intrusion in Associative Memory and LLMs
Abstract
A memory system balances two competing demands: recovering stored patterns from partial cues and keeping similar experiences distinct. We study one failure of this balance, \emph{source intrusion}, in which content that is valid in one context is selected when another source is queried. We conduct two complementary studies. First, a controllable associative-memory model compares fixed retrieval with familiarity-dependent suppression, allowing us to test whether retrieval strength should adapt to cue state. Second, a controlled two-report task induces source conflict in three instruction-tuned large language models (LLMs) and probes their pre-generation hidden states for future errors. Familiarity-dependent control improves lure discrimination across three levels of cue degradation. In the LLMs, error information is most decodable before the final representation, and intermediate-layer gates reduce held-out risk relative to final-layer gates at matched realized coverage. These results connect adaptive associative retrieval with source-aware monitoring and suggest that memory reliability depends not only on what information is available, but also on how competing memories are controlled and where their conflicts are monitored.