POST: Progressive Object-Slot Tokenization for Multimodal Large Language Models
Abstract
Multimodal large language models (MLLMs) leverage object-agnostic tokenization by default, which represents images as flat grids of patch tokens and uniformly discretizes semantically rich foreground regions and information-sparse background areas. This introduces substantial redundancy and constitutes a major source of inference cost. Existing object-level tokenizers could either discard intra-region saliency with predefined region grouping and coarse pooling or limit to a single semantic scale by employing slot attention to the final encoder feature map only. In this paper, we propose a two-stage framework named POST that decouples unsupervised object-centric grouping from language-grounded multimodal reasoning to achieve progressively semantic-aligned object-level visual tokenizers. Specifically, we develop Progressive Slot Attention (PSA) in Stage I to introduce slot attention across multiple layers of a frozen ViT encoder and propagates slot states through a cross-layer refinement chain for learning hierarchical slot--patch correspondences without segmentation supervision. Subsequently, we transfer the learned PSA to the LLaVA pipeline in Stage II for object-level visual tokenization. We design slot-weighted merging to aggregate ViT patch tokens into object-grounded visual tokens and remove low-support slots with adaptive slot pruning. Remarkably, POST eliminates the need for external segmentation models, enables progressive object-centric refinement, and adapts visual token budget to image complexity in visual tokenization. Experimental results show that PSA achieves strong unsupervised object-centric segmentation on PASCAL VOC and COCO. Furthermore, POST is comparable or superior to LLaVA-1.5 on most multimodal benchmarks using only 8.3% visual token budget and improves referring expression comprehension on RefCOCO/+/g. It also outperforms recent patch-level and object-level token compression baselines at comparable budgets.