Transporting Quantiles to Codebook: Scalable Vector Quantization without Codebook Collapse
Yuwei Zeng ⋅ Zekun Shi
Abstract
Effective discrete representation learning with VQ-VAE serves as a fundamental component of modern autoregressive generative models. However, it is widely known to suffer from codebook collapse and inefficient code utilization. To address this issue, we propose a simple yet effective alternative to standard vector quantization by formulating neural quantization as a joint modeling and clustering process. Specifically, we learn the latent space using a lightweight flow-matching model alongside training, corresponding to a probability flow ODE that transports an often fixed well-behaved source distribution while preserving probability mass. Based on this, we define the codebook $\mathcal{C}_1$ by transporting even-quantile partitioned prior codes $\mathcal{C}_0$ from the source space through the learned flow, and obtaining the Voronoi centroids under the induced distribution, leading a structured discretization. We evaluated on image tasks, and two tasks require high-precision tokenization, demonstrating scalability and improved expressivity of the codebook with full utilization.
Chat is not available.
Successful Page Load