PixelDiT2: Representation-Grounded Pixel Diffusion Transformers
Yongsheng Yu ⋅ Wei Xiong ⋅ Yichen Sheng ⋅ Shiqiu Liu ⋅ Jiebo Luo
Abstract
Recent advances in pixel-space diffusion models have narrowed the image quality gap with latent-space diffusion, but still converge more slowly and lag behind in final image quality. We argue that a key reason is the lack of an explicit representation prior: unlike latent diffusion, which denoises in a compact and structured latent space, pixel diffusion must learn denoising-friendly representations and pixel generation simultaneously from raw RGB space. To address this problem, we propose PixelDiT2, an end-to-end pixel-space diffusion model designed to decouple representation learning from pixel generation without introducing an autoencoder or latent reconstruction bottleneck. We propose *representation grounding* that uses a frozen pretrained vision foundation model to provide explicit per-patch representation guidance throughout denoising, allowing the pixel diffusion transformer to focus more on pixel generation. On ImageNet-$256{\times}256$, PixelDiT2 reaches FID $\textbf{1.50}$ in $320$ epochs, surpassing JiT-G's FID $1.82$ at $600$ epochs with roughly half the parameters and half the training budget.
Chat is not available.
Successful Page Load