Where to Look Is Not How to Fix: Pre-Denoising Diagnostics and Modality-Dependent Control in Diffusion Composition
Fangzheng Wu ⋅ Brian Summa
Abstract
Text-to-image diffusion models fail on semantically rare attribute--object compositions, but it is unclear whether this brittleness arises during denoising or is already present in conditioning representations. We study this question at three levels of granularity. First, using a controlled stress protocol across SD1.5, SDXL, and SD3, we show that a Compositional Stress Index derived from text-side non-factorization residuals separates anchor from stress prompts with perfect discrimination (AUC${=}1.0$) cross-model, establishing a robust pre-denoising risk axis. Second, embedding-level low-rank repair reduces representation residuals without any positive semantic improvement, while a downstream cross-attention intervention succeeds on the same subset---a site-sensitive boundary between diagnosis and control. Third, we decompose this boundary at block-group resolution under three intervention modalities. Diagnostic accessibility peaks at the deep encoder (D2) across all conditions, but token-level intervention of either sign produces no measurable improvement at any block. Only broad cross-attention ablation moves semantics at D2, while decoder blocks respond to selective boost and collapse under ablation. Together, these results establish that compositional defects are diagnosable before denoising, but \textbf{where to look is not how to fix}: the diagnostic site and the control site require different intervention modalities, and even at the correct site, the wrong modality is futile.
Chat is not available.
Successful Page Load