Do Adaptive Image Tokenizers Spend Tokens Where Vision Needs Them
Abstract
Adaptive image tokenizers promise to vary representation length with visual con tent, but their quality is still commonly trained, selected, or reported through image reconstruction. We ask a narrower question: when token capacity is scarce, does marginal reconstruction improvement identify where tokens are useful to down stream vision? We audit four heterogeneous adaptive or flexible-rate tokenizers— KARL, ALIT, ElasticTok, and FlexTok—on 4,952 COCO images, five native rates per tokenizer, and three frozen tasks: caption retrieval, object recognition, and semantic segmentation. Per-image LPIPS gain has only weak rank correlation with downstream gain (ρ = 0.029–0.132). At identical total token costs, allocating rates using measured task gains outperforms allocating them using measured reconstruc tion gains at all 36 interior tokenizer–task–rate operating points; the result holds for LPIPS, L1, SSIM, and PSNR (144/144 comparisons). A controlled spatial restoration audit on 576 frozen cases localizes the mismatch: reconstruction- and task-ranked regions are nearly uncorrelated, task-ranked restoration wins every comparison at 12.5–50% restored area, and task benefit concentrates more strongly on small objects and boundaries. Reconstruction remains meaningful for fidelity, but is incomplete as an allocation signal. Our results motivate evaluating flexi ble tokenizers by matched-rate downstream allocation, rather than rate–distortion curves alone