ShadowBench: Exposing Lexical Anchoring and the Illusion of Forgetting in Large Language Models
Abstract
Large Language Models (LLMs) have emerged as primary interfaces for factual retrieval, yet current evaluation paradigms rely almost exclusively on explicit entity names. In this work, we demonstrate that LLM knowledge is fundamentally lexically anchored: models possess extensive factual information but find it significantly difficult to retrieve without explicit name tokens. To quantify this limitation, we introduce ShadowBench, a rigorously hardened, shortcut-resistant benchmark designed to evaluate Latent Entity Association via a novel Dual-Trait Association (DTA) task. Our evaluation reveals a pervasive "Shadow Gap" across all model scales – up to the frontier models GPT-5.4 and Claude-Sonnet-4.6 – where removing lexical anchors causes performance drops of over 20%. Furthermore, we demonstrate that this lexical dependency exposes a critical vulnerability in AI safety. Applying ShadowBench to Machine Unlearning, we find that state-of-the-art algorithms (e.g., Gradient Difference, NPO) achieve only superficial lexical erasure. While unlearned models fail on direct queries, our novel Latent Entity Leakage Rate (LELR) metric reveals that reasoning models explicitly reconstruct the "forgotten" entity in their internal reasoning traces in over 89% of cases, utilizing residual shadow knowledge to solve associative tasks. Ultimately, ShadowBench proves that current unlearning paradigms create an "Illusion of Forgetting," demonstrating that existing methods act as superficial output filters and are not yet reliable for achieving true parametric erasure.