Memory Busted: Activation-Based Provenance Catches LLMs Cheating on Their Context
Abstract
Large Language Models (LLMs) frequently generate responses that deviate from the external context provided to them, undermining reliability even within Retrieval-Augmented Generation (RAG) pipelines. This failure arises from a conflict between the model's parametric memory and the retrieved context at inference time; when memory wins, the response is often fluent and confident yet ungrounded, a dangerous failure mode in high-stakes domains such as medicine and finance, where outputs must be traceable to their source. We show that this conflict leaves a distinct, recoverable signature in the model's feedforward activations, and introduce Activation-based Response Provenance (ARP), a recurrent probing classifier that models the temporal trajectory of token-level activations to determine whether a generated response is grounded in the provided context or drawn from memory. Across five QA datasets under controlled and random perturbations, ARP consistently outperforms existing baselines, and remains robust in a black-box setting using a different off-the-shelf LLM for activation extraction. These results establish feedforward activation trajectories as a robust, model-agnostic signal for auditing whether LLM generations remain faithful to context or are silently overridden by parametric memory.