Mitigating Memorization Where It Happens
Xuanqi Zhang ⋅ HaoYang Shang ⋅ Xiaoxiao Li
Abstract
Large language models (LLMs) can memorize and reproduce training sequences verbatim, which undermines both generalization and privacy. Existing mitigation methods apply interventions uniformly, often degrading performance on the majority of tokens that generalize normally. We empirically find that memorization is sparse and heavy-tailed at the token level. This structure is mismatched to static, sequence-level interventions: effective mitigation must localize to where memorization actually happens. From the ideal token-level memorization reduction objective, we derive a probe-steer framework, which decomposes intervention into a probe that detects memorization-relevant activations and a steer that applies targeted correction only when the probe exceeds a threshold. We propose Gated Subspace Steering (GSS) as its practical instantiation: the optimal probe-steer pair admits a closed-form solution as the leading singular vectors of a gradient-weighted activation matrix. Across four benchmarks, GSS matches or exceeds state-of-the-art memorization reduction while requiring 10--100$\times$ less compute than optimization-based alternatives.
Chat is not available.
Successful Page Load