PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion
Yifan Lu ⋅ Qi Wu ⋅ Jay Zhangjie Wu ⋅ Zian Wang ⋅ Huan Ling ⋅ Sanja Fidler ⋅ Xuanchi Ren
Abstract
Most practical high-resolution text-to-image systems rely on latent diffusion models, where generation is performed in a compact latent space and a decoder maps latents back to pixels. Yet the latent-to-pixel decoder in this pipeline remains primarily reconstruction-oriented: as the sole pathway from latents to pixels, it is optimized to invert an encoder rather than to synthesize high-resolution details, and becomes increasingly costly in megapixel pipelines. This gap calls for a more expressive and efficient decoding paradigm. Motivated by recent progress in scalable pixel-space diffusion, we introduce **PiD**, a **Pi**xel diffusion **D**ecoder that reformulates latent decoding as conditional pixel diffusion, unifying decoding and upsampling into one generative module. By denoising directly in high-resolution pixel space, PiD synthesizes $4\times$and even $8\times$ upscaled images with low latency. For latent conditioning, a lightweight sigma-aware adapter injects noise-corrupted latents into the backbone, enabling PiD to handle partially denoised states and terminate the base diffusion process early. To further improve efficiency, we distill the decoder using DMD2, reducing inference to just $4$ steps. PiD applies to both conventional VAE latents and semantic latents (e.g., SigLIP, DINOv2) used in recent RAE-based models. Across multiple latent spaces and base generators, PiD improves visual fidelity over conventional decode-then-upsample cascades while reducing memory and latency, decoding latents of $512{\times}512$ images into $2048{\times}2048$ pixels in under $1$ second with $13$ GB peak memory on a consumer RTX 5090, and as fast as $210$ ms on a data-center GPU, about $6{\times}$ faster than cascaded diffusion-based super-resolution pipelines.
Chat is not available.
Successful Page Load