Pixel-space Autoregressive Image Synthesis via Spectrum Serialization and Flow-based Refinement
Abstract
Autoregressive (AR) image generators commonly rely on vector-quantized (VQ) autoencoders to compress images into discrete token sequences. However, quantization inevitably discards visual information, and the resulting reconstruction errors may propagate through AR decoding, limiting generation fidelity. We revisit the representation choice for AR image generation and ask whether a simpler and more faithful representation can alleviate this bottleneck while remaining compatible with autoregressive modeling. To this end, we estimate the intrinsic dimension of the natural image manifold under several widely used representations and find that raw patchified pixels exhibit the simplest underlying geometry among those evaluated. Motivated by this insight, we propose PixeLLM, an AR generative framework that operates directly on sequences of discrete pixel blocks, eliminating the need for a separately trained VQ autoencoder. PixeLLM decomposes generation into two stages: (1) an AR model generates a text-aligned draft image capturing global scene structure, and (2) a conditional flow-based pixel refinement Transformer enhances the draft with fine-grained details and mitigates artifacts. The resulting pipeline is straightforward, requiring no external semantic alignment. Despite its simplicity, PixeLLM achieves competitive performance on challenging text-conditioned and class-conditioned image generation benchmarks.