TrackTok: Object-Centric Video Tokenization with Semantically Persistent Tokens
Abstract
We introduce TrackTok, an efficiency-focused video tokenization framework that produces object-centric tokens, i.e., tokens that consistently encode the same semantic objects across time, even as it moves, deforms, or undergoes partial occlusion. Conventional patch-based or existing holistic tokenizers produce token identities often tied to fixed image locations. In contrast, our approach learns temporally persistent object-centric representations that bind visual features across frames into stable semantic units. This enables substantially more compact, continuous latent video representations, reducing the number of tokens required for downstream generative modeling with state-of-the-art flow models, while preserving the semantic structure needed for high-quality synthesis. On the UCF-101 video generation benchmark, we provide a proof of principle that our tokenizer uses ~4x fewer latent tokens to achieve the same generative FVD compared to existing holistic tokenizers. This results in an up to 10x increase in sample throughput, offering a favorable quality-efficiency trade-off. These results suggest that object-aligned and temporally persistent tokenization is a promising direction for scalable video generation, where reducing latent token counts is critical for efficient modeling.