The Illusion of Forgetting: Rank Leakage in Knowledge Editing and Its Mitigation
Jie Hou ⋅ Qiang Zeng
Abstract
Knowledge editing aims to rewrite sensitive information in language models, yet prior work has shown that edited models remain vulnerable to extraction attacks. In this work, we identify a previously overlooked vulnerability, which we term $\textit{rank leakage}$: although editing replaces the top-1 output with a sanitized answer, the original sensitive token often remains highly ranked in the model’s next-token distribution. Exploiting this phenomenon, we propose an effective black-box Output Attack that recovers sensitive information using only a single query. Despite its simplicity, the attack achieves strong extraction performance and remains effective even against models equipped with state-of-the-art defenses. To mitigate this vulnerability, we introduce $\textbf{R}$ank-guided $\textbf{C}$oncealment $\textbf{E}$diting (RCE), a plug-in defense that enforces explicit control over the rank of sensitive tokens during editing. Although designed for the Output Attack, RCE generalizes to existing black-box and white-box extraction attacks. Extensive experiments across multiple models, datasets, and editing algorithms show that RCE consistently improves resistance to attacks while preserving edit fidelity and model utility. Our results highlight the importance of controlling the rank of sensitive tokens, rather than only the top-1 prediction, for privacy-preserving knowledge editing.
Chat is not available.
Successful Page Load