Representation Preconditioning for Efficient Diffusion Model Training
deyuan liu ⋅ PENG SUN ⋅ Xufeng Li ⋅ Tao Lin
Abstract
Diffusion transformers learn useful internal representations during training, and learning them from scratch can add a burden to efficient training. Alignment mitigates this burden by adding an auxiliary loss that regularizes diffusion hidden states with features from pretrained visual encoders. In this work, we revisit the role of this alignment signal. Rather than using pretrained representations only as a regularizer during full diffusion training, we use them as a preconditioning signal before diffusion training. The key observation is that early layer alignment from clean latents to pretrained representations can be much cheaper than optimizing the full flow matching objective over the entire backbone, yet provides a better initialization for subsequent diffusion training. We instantiate this idea as Embedded Representation Warmup (ERW). ERW first aligns early layers from clean latents to pretrained representations, and then trains the full model with the standard diffusion objective and decaying alignment. This initialization allows subsequent diffusion training with alignment to converge more efficiently than starting alignment from scratch. Counting warmup in the total budget, ERW improves convergence in SiT based diffusion training. On ImageNet $256\times256$, ERW reaches FID=1.41 in 350 epochs with SiT-XL/2; it also reaches FID=2.04 on ImageNet $512\times512$ and improves FID on MS-COCO text-to-image generation.
Chat is not available.
Successful Page Load