You Can’t Have It Both Ways: Concept Entanglement Limits Diffusion Model Unlearning
Abstract
Concept unlearning in text-to-image diffusion models aims to suppress a target concept, e.g., \texttt{horse}, while preserving the model's ability to generate semantically related but distinct content, like \texttt{donkey}. Yet existing methods either leak under indirect prompts or visibly degrade remaining concepts. {\em We show that these failure modes arise naturally from overlapping concept representations.} Our main contribution is demonstrating that the observed trade-offs stem from the geometry of concept representations, rather than from weaknesses in any particular algorithm. We support this claim through both theoretical analysis and empirical observations. By formalizing concepts as activation-space regions, we show that overlap between a target and other concepts lower-bounds the unavoidable degradation on those other concepts when the target is erased. Our analysis suggests that strong erasure and preservation become fundamentally coupled when concepts occupy overlapping activation regions, with the trade-off scaling linearly in the degree of overlap. We verify this trade-off on seven unlearning methods spanning fine-tuning, adversarially-robust fine-tuning, and inference-time interventions. Averaged across target concepts, STEREO almost completely suppresses the target concept under indirect prompts but cuts the model's ability to generate semantically related concepts by more than 75\%. On the other hand, sparse inference-time methods (SAeUron, SEOT) better preserve utility but leave substantial target leakage. The trade-off is steep for concepts whose internal representations heavily overlap with their evaluated semantic neighborhoods, (e.g., \texttt{horse} with \texttt{cat}, \texttt{dog}, and \texttt{bear}) and milder for \texttt{castle}, which is comparatively less entangled within our concept pool and probe layer in activation space. Across targets, damage to related concepts scales monotonically with our activation-overlap measure. These results suggest that perfect unlearning is the wrong target for entangled concepts. The field should evaluate on the Pareto frontier our theorem establishes; current benchmarks, which decouple unlearning accuracy from utility preservation, hide this trade-off and need to be revised.