RankE: End-to-End Post-Training for Discrete Text-to-Image Generation with Decoder Co-Evolution
Abstract
Discrete autoregressive (AR) text-to-image (T2I) models adopt a two-stage paradigm in which a VQ tokenizer maps images to discrete codes and an AR policy models their distribution. Current post-training methods optimize only the AR policy while keeping the VQ decoder frozen. We show that this practice introduces Latent Covariate Shift: as the policy evolves, its token distribution progressively diverges from the ground-truth distribution on which the decoder was trained, such that reward scores improve while decoded image quality degrades. To address this mismatch, we propose RankE, the first end-to-end post-training framework for discrete T2I generation. Rather than optimizing the policy against a fixed decoder, RankE co-evolves the policy and decoder through alternating optimization: the policy is refined via group-relative preference optimization, while the decoder is jointly adapted via reward-aware adversarial training. This co-evolution ends the fidelity--alignment trade-off that plagues frozen-decoder approaches: on LlamaGen-XL (775M), standard RL improves CLIP but degrades FID, whereas RankE simultaneously improves both (FID 15.21, CLIP 33.76 on MS-COCO 30K). Consistent joint gains on Janus-Pro (1B) confirm that decoder co-evolution reliably converts reward optimization into pixel-space quality improvements.