Parallel Rollout Approximation for Pixel-Space Autoregressive Image Generation
Jiayi Xu ⋅ Di He ⋅ Guolin Ke
Abstract
Continuous-token autoregressive (AR) generation directly models continuous signals without discretizing them into vocabulary tokens, but suffers from a severe train--inference mismatch: teacher-forced training conditions on ground-truth prefixes, whereas inference conditions on model-generated continuous tokens. This mismatch becomes especially severe in pixel-space image generation, where each token is a high-dimensional raw pixel patch and autoregressive errors can accumulate over generation steps. Rollout-based training can reduce this mismatch but is prohibitively slow, especially when each token is produced by a multi-step diffusion head. We propose \emph{Parallel Rollout Approximation} (PRA), which approximates rollout-based training by constructing training inputs in parallel at all positions. PRA combines a shared one-step denoiser with an end-to-end learned low-dimensional intermediate state, aligning these constructed training inputs with inference-time generated outputs without relying on a separately pretrained tokenizer. On class-conditional ImageNet-1K generation at 256$\times$256 resolution, even PRA-S with 135M parameters achieves an FID of 2.58, surpassing the previous billion-scale pixel-space AR result of 3.60. Scaling to PRA-L (511M) further improves FID to 1.94, establishing a new state of the art among pixel-space AR models and substantially narrowing the gap to pixel-space diffusion models. Beyond generation, PRA achieves higher ImageNet classification probing accuracy than latent-space AR baselines, suggesting its potential for unified pixel-space image generation and understanding.
Chat is not available.
Successful Page Load