SCULPT: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis.
Abstract
We propose SCULPT, a state-of-the Art masked discrete diffusion model (MDMs) for high resolution text-to-image synthesis. Compared with prior works on masked image generation, SCULPT addresses two key challenges. First, unlike continuously diffusion models which progressively refines image latent across the entire image, vanilla MDMs do not have self-correcting capability because discrete tokens cannot be changed once it's unmasked. Second, while scaling the vocabulary size of discrete image tokenizers can improve reconstruction quality, it also introduces optimization challenges for training generative models as per-token training signal becomes more spares. To address the first challenge, SCULPT incroprates a token-editing mechanism where the model can dynamically correct already-unmasked output tokens during inference. To address the second challenge. we proposed a Goruped Corss Entrophy (GCE) objective that assigns positive learning signal to adjacent tokens of the ground truth in the embedding space. To improve the training efficiency, we further implemented a custom fused operator that greatly reduce the VRAM requirements during training in large-vocabulary setting. Experiment results show that these innovations significant improves the training efficiency and image fidelity of masked discrete image generators.