EntiRE: Invariant Learning for Robust Concept Erasure in Text-to-Image Generative Models
Abstract
Large-scale text-to-image diffusion models can inadvertently generate harmful or sensitive content learned from uncurated training data, motivating the development of concept erasure methods that remove undesired concepts from pre-trained models. However, recent studies reveal that erased models remain vulnerable to adversarial prompt attacks that recover the supposedly removed concepts, highlighting the need for robust erasure techniques. In this work, we visualize the text embedding distributions of adversarial prompts and find that the difficulty of erasure is non-uniform. Existing methods effectively erase concepts in easier regions but leave localized residual holes in harder regions, which adversarial prompts consistently exploit. Motivated by this finding, we propose Entire-space Robust Erasure (EntiRE), an end-to-end framework that casts robust concept erasure as an out-of-distribution generalization problem. By adopting an invariant learning formulation regularized by total variation, EntiRE enforces more uniform erasure across the continuous prompt space, suppressing the localized residual regions. Extensive experiments on erasing nudity, artistic style, and object concepts demonstrate that EntiRE consistently outperforms state-of-the-art baselines in both robustness against adversarial prompt attacks and utility preservation.