Memorized or Memo-rized: Prior Leakage of Procedural Knowledge in Large Language Models
Kexin Wang ⋅ Lin Li ⋅ Yi Ren ⋅ Zihuiwen Ye ⋅ Gusheng Pan ⋅ Yug D Oswal ⋅ Yarin Gal ⋅ Yihong Chen
Abstract
Agentic AI systems combine foundation models’ internal knowledge with external knowledge from retrieval, tools, or memory modules. Conventionally, these are entangled with reasoning ability, making it hard to isolate how a model generalizes under internal--external knowledge conflict. In this paper, we formalize knowledge into $\textit{abstract}$ and $\textit{surface-form}$ knowledge, and $\textit{reasoning}$ as a way to localize the effects of knowledge conflict. With this, we identify the $\textbf{prior leakage}$ phenomenon: a model’s over-reliance on internal knowledge under internal--external conflict. We evaluate 17 open source models on two newly proposed procedural knowledge datasets, $\textbf{SQLShift}$ and $\textbf{PandasShift}$, where common SQL and Pandas operators are swapped or renamed. We observe extensive prior leakage across the models: knowledge stays heavily anchored on internal knowledge, even when explicitly instructed to use external knowledge. As an ablation, we applied gradient ascent unlearning to forget prior knowledge, which did not substantially change prior leakage. We also attempted to control internal knowledge by post-training a vintage LLM, through varying number of training epochs and size of training set. We believe this evaluation suite will support the development of trustworthy AI agents that respect user-provided external evidence over unwarranted reliance on parametric priors, paving the way toward user-sovereign AI.
Chat is not available.
Successful Page Load