Contrastive Representation Learning for Robust Concept Erasure in Diffusion Models
Samuel Simko ⋅ Rishit Dagli ⋅ Zhijing Jin ⋅ Bernhard Schölkopf
Abstract
Concept erasure edits a pretrained text-to-image diffusion model to remove an unwanted or unsafe concept while preserving its other capabilities. Existing methods remain vulnerable to adversarial prompts and optimized conditioning embeddings, showing that suppression does not ensure robust erasure. We introduce CIRCE (Contrastive Image-Representation Concept Erasure), a contrastive objective on the denoiser's internal representations. It moves target representations away from their values in the frozen model and toward a shared batch anchor while preserving benign representations. CIRCE sets a new best mean on UnlearnCanvas among ten published defenses. For nudity erasure on Stable Diffusion 1.4, our method achieves the lowest worst-case attack success across seven attacks. Its largest advantage is under continuous concept inversion, where it reaches $0.138$ compared with $0.225$ for the next-best method, while ranking second among erased models in retain-quality Elo. Ablations show that this robustness comes primarily from grounding the forget set in images of the concept rather than strengthening a token-only collapse. These results establish representation-level intervention as an effective route to robust concept erasure when the forget data cover the target capability.
Chat is not available.
Successful Page Load