Rethinking Cross-Layer Information Routing in Diffusion Transformer
Chao Xu ⋅ Maohua Li ⋅ Qirui Li ⋅ Yixuan Xu ⋅ Yanke Zhou ⋅ Yunhe LI ⋅ Cuifeng Shen ⋅ Hanlin Tang ⋅ Kan Liu ⋅ Tao Lan ⋅ Lin Qu ⋅ Shao-Qun Zhang
Abstract
Diffusion Transformers (DiTs) have become the de facto backbone of modern visual generation, and nearly every major axis of their design --- tokenization, attention, conditioning, objectives, and latent autoencoders --- has been thoroughly revisited. The residual stream that governs how information accumulates across layers, however, has been directly inherited from the original Transformer. In this paper, we present a systematic analysis of cross-layer information flow in DiT jointly along depth and denoising timestep, and identify three concrete symptoms of traditional residual addition, i.e., monotonic forward magnitude inflation, sharp backward gradient decay, and pronounced block-wise redundancy. Motivated by this diagnosis, we propose Diffusion-Adaptive Routing (\textsc{DAR}), a drop-in residual replacement that performs \emph{learnable, timestep-adaptive, and non-incremental} aggregation over the history of sublayer outputs. Moreover, the proposed DAR is compatible with many modern Transformer enhancement methods, such as REPA. On ImageNet $256\times256$, \textsc{DAR} improves SiT-XL/2 by $2.11$ FID ($7.56$ vs.\ $9.67$) and matches the baseline's converged quality in $8.75\times$ fewer training iterations. Stacked on top of REPA, it yields a $2\times$ training acceleration in the early stage, revealing cross-layer information routing as a hitherto-overlooked axis of progress in diffusion modeling, one that operates orthogonally to existing representation-alignment objectives. Beyond the pretrain tasks, \textsc{DAR} can be further applied in large-scale T2I models fine-tuning stage and preserves high-frequency details during Distribution Matching Distillation.
Chat is not available.
Successful Page Load