Defense-in-Depth for LLMs: Evaluating Memory Gates Against Activation-Induced and Memory-Induced Sycophancy
Abstract
Long-term memory allows Large Language Models (LLMs) to maintain personalized context across interactions, but retrieved user history can induce memory-induced sycophancy, causing models to favor stored user beliefs over objective evidence. Existing defenses primarily operate on retrieved context and are rarely evaluated jointly with internal behavioral bias. We introduce a 2×2 defense-in-depth framework separating internal activation steering from external memory handling. We extract sycophancy steering directions from 100 paired prompts and evaluate four open-weight models across 10 steering coefficients, five memory-defense configurations, and all 1,550 MemSyco-Bench items, with three LLM judges. (Footnote: Our anonymized code, generation reports, and judge logs are available at https://anonymous.4open.science/r/ourprojectsubmissionpt1/ and https://anonymous.4open.science/r/ourprojectsubmissionp2/.) Selective Router Gate filtering preserves substantially more of MemSyco's average accuracy than complete memory removal, and this separation persists as internal sycophancy pressure is raised. On Llama 3.1 8B, mild inverse steering with Router Gate reduces judge-averaged sycophancy from 35.80% to 31.32%, while average accuracy decreases from 43.99% to 43.31%. These results show that external memory filtering provides the robust layer of a defense-in-depth design, while internal inverse steering adds a smaller further reduction with a measurable safety–utility trade-off.