One Carrier Isn't Enough: Redundancy-Aware Concept Suppression in Diffusion Transformers
Abstract
Concept suppression, the ability to prevent a text-to-image model from generating a specific object or style, is a core requirement for deploying generative models safely: blocking harmful or infringing content and giving operators a lightweight tool to intervene without retraining. Existing approaches typically rely on weight editing or retraining, which is costly and hard to update. We study a training-free alternative for modern Diffusion Transformers (DiTs): editing the text-conditioning representations a pretrained model already computes. DiTs route text through multiple architecture-specific carriers, for example a pooled embedding driving adaptive normalization and one or more sequence embeddings entering joint attention. We hypothesize that concepts are redundantly represented across these carriers, making single-carrier erasure structurally insufficient. We introduce joint carrier erasure, a training-free framework that identifies each model's conditioning carriers, learns a low-rank concept subspace per carrier from contrastive prompt pairs, and jointly projects the concept out of every carrier at generation time. On a structured style by object benchmark across FLUX-schnell, FLUX-dev, and SD3.5-medium, joint carrier erasure achieves strong object suppression while preserving non-target content. We introduce a retention-aware evaluation protocol, including a four-way audit of removal outcomes, to expose failure modes such as image destruction that simpler accuracy metrics can mask. We also find that style suppression proves substantially harder and more model-dependent. Carrier ablations confirm single-carrier edits are unreliable compared to joint erasure, supporting the view that concept information in DiTs is distributed across the conditioning pathway rather than localized to one representation.