Where and When Identity Forms: Identity-Vital Attention Redistribution for Training-Free Subject-Driven Generation
Luan Thanh Trinh ⋅ Atsuki Osanai ⋅ Kenji Doi
Abstract
Subject-driven diffusion transformers preserve coarse object identity but routinely lose identity-defining details such as logos, printed text, textures, and small structural cues. We show that this failure traces to a structural mismatch: identity sensitivity is sparse and non-uniform across layers, denoising stages, and spatial regions, yet reference influence is typically applied as a global signal. We introduce \textbf{VitalScore}, a training-free method that discovers this identity-vital structure through offline probing and uses it to build a reference-attention redistribution tensor $\Lambda(l,t,s)$ coordinated with a stage-aware prompt schedule $(P(t)$. The spatial axis is grounded in a two-phase VLM blueprint calibrated against human identity annotations, ensuring the controller reflects the cues people actually use to judge subject identity. We evaluate across Diptych Prompting, FLUX Kontext, and FLUX 2.0 on DreamBench++, where VitalScore improves DINO-v2 by up to 15%, and on a new product identity benchmark targeting fine-grained commercial cues that standard benchmarks overlook, where gains reach 16.5% on the hardest prompt axis---all without parameter updates.
Chat is not available.
Successful Page Load