RefineTok: Scale-Wise Tokenization for Progressive Visual Refinement
Abstract
Image tokenization aims to compress visual signals into latent codes that remain effective for reconstruction and generation. Driven by the multi-scale nature of images, organizing latent along a coarse-to-fine refinement process emerges as a promising direction for visual tokenization. Existing scale-wise tokenizers often introduce this structure through token ordering or latent prefixes, but each target scale is typically decoded directly from its prefix, making different prefixes behave like separate reconstruction codes. This can lead to redundant encoding of global structure across scales and limits the effectiveness of compact token budgets. We propose RefineTok, a scale-wise continuous tokenizer that makes decoding itself progressive. RefineTok carries a visual state across scales and uses each latent prefix to refine the previous state into the next scale, turning scale-wise latents into incremental refinement signals rather than independent reconstruction codes. This design yields a compact and structured latent interface for both reconstruction and downstream generation. Together with prefix-level training, a coarse-centric scale trajectory, and a semantic anchor, RefineTok achieves strong reconstruction fidelity among scale-wise tokenizers and competitive class-conditional generation on ImageNet-1K under compact latent budgets. Meanwhile, it preserves a progressive scale-wise interface that single-scale tokenizers do not provide.