SpaG-DiT: Enhancing Spatial Grounding for Diffusion Transformers
Abstract
Despite remarkable progress in text-to-image generation, current Diffusion Transformers (DiTs) inherently struggle with learning intricate visual content, such as accurate multilingual scene text rendering and novel visual concepts. This limitation stems from the inherent contradiction between vague semantic guidance provided by text and the objects of accurate reconstruction during diffusion training. To address this bottleneck, we propose SpaG-DiT, a novel spatial grounding training framework. SpaG-DiT directly injects explicit spatial priors into the DiTs before the attention module via Contextual Diagonal Position Encoding (CDPE), which dynamically modulates spatial coordinates without destroying the inherent sequential structure of the linguistic input. Furthermore, we introduce VLM Grounding Guidance (VLM-GG) as an automated pipeline to extract implicit spatial priors from pre-trained Vision-Language Models, eliminating human annotation costs. To evaluate the grounding ability, we establish two new datasets: SpaG-DiT-MultiLingual for non-Latin/Chinese text rendering and SpaG-DiT-Creature for new concept learning. Extensive experiments demonstrate that SpaG-DiT achieves superior training efficiency and significantly outperforms vanilla self-finetuning on the above two tasks and the widely adopted GenEval benchmark. These results validate the plug-and-play nature of SpaG-DiT and its potential to enhance DiT training processes. Code, model, and data will be released.