Phase-wise Velocity Distillation: Towards Effective Image Generation with A Single NFE
Zhen Guo ⋅ Rongyuan Wu ⋅ Qiaosi Yi ⋅ Chenxi Xie ⋅ Xinyu Wei ⋅ Lei Zhang
Abstract
While diffusion distillation methods have largely accelerated image generation, achieving high-quality synthesis under a single number of function evaluations (NFE) budget remains a challenging problem. The diffusion process follows a coarse-to-fine progression, yet existing solutions typically force a single student model to simultaneously resolve global structures and fine details in one forward pass, leading to over-smoothed outputs. To address this issue, we propose **P**hase-wise **V**elocity **D**istillation (**PVD**), which strategically partitions the generation timeline into a coarse and a fine phase, and models the transition within each phase via the average velocity. A dedicated half-sized expert is then assigned to each phase, decoupling structural composition from detail refinement while keeping the cumulative cost within a single-NFE budget of the teacher. While this design already suffices for class-conditional image (C2I) generation, we further introduce phase-wise adversarial supervision with dedicated discriminators for the more complex text-to-image (T2I) tasks, ensuring accurate distribution matching. On C2I generation, PVD achieves a state-of-the-art FID of 1.48 under a single-NFE budget on ImageNet $256 \times 256$. On T2I tasks, PVD-distilled models (Stable Diffusion 3.5-Medium, FLUX.1-dev) produce results competitive with their multi-step teachers, significantly outperforming prior distillation methods. Source codes and distilled models will be released.
Chat is not available.
Successful Page Load