Structural Self-Teaching for Compositional Generation in Unified Multimodal Models
Abstract
We propose Structural Self-Teaching (SST), a novel framework that enhances compositional image generation in unified multimodal models (UMMs) by leveraging their internal understanding path as a source of structural supervision. The motivation of this work is the observation that the generation path of UMMs often struggles with dense compositional prompts, leading to missing objects, attribute binding failures, and spatial errors. To bridge this gap, our key idea is to convert the phrase-level grounding inherent in the understanding path into training-time structural signals. Specifically, we extract phrase-conditioned support masks from internal attention maps derived from the understanding path and employ them through two complementary mechanisms: 1) phrase-local contrastive alignment to synchronize generated features with their corresponding phrase-level grounding, and 2) a lightweight structural scaffold for spatial modulation of intermediate generation states. By employing attenuated scaffold dropout during training and removing the scaffold branch at inference, our approach improves compositional fidelity across object, attribute, and spatial constraints. Crucially, the framework requires no external structural annotations for training or additional inputs at test-time, preserving the original generation interface. Extensive experiments across multiple UMMs demonstrate that our method consistently improves compositional fidelity while maintaining the efficiency of a lightweight tuning setup.