OneVision-Encoder: Codec-Aligned Sparsity as a Foundational Principle for Multimodal Intelligence
Abstract
Visual signals, especially videos, are highly redundant: most regions are temporally predictable, while informative changes are sparse. We hypothesize that vision encoders should allocate computation according to this non-uniform information density rather than process dense pixel grids uniformly. From this perspective, codec-derived motion and residual signals provide a natural basis for identifying informative regions. OneVision-Encoder instantiates this idea through Codec Patchification, which replaces uniform dense computation with selective processing of only 3.1\%-25\% of regions that carry high signal entropy. To support irregular spatiotemporal token layouts, OneVision-Encoder employs a shared 3D RoPE and is pretrained with a large-scale cluster discrimination objective over more than one million semantic concepts, enabling unified representation learning for appearance and motion. Empirically, OneVision-Encoder improves both efficiency and accuracy under fixed token budgets. When integrated into large multimodal models, it achieves a 4.1\% average improvement over Qwen3-ViT on video understanding benchmarks while remaining competitive on image and document understanding. Under attentive probing, it further yields substantial gains on motion-sensitive benchmarks, improving Top-1 accuracy on Diving48 by 17.1 and 8.1 percentage points over SigLIP2 and DINOv3, respectively, at matched patch budgets. These results suggest that codec-guided patch sparsity is an effective design principle for scalable visual representation learning.