Video Token Compression via Formation and Selection Distillation in Large Vision Language Models
Soowon Son ⋅ Heeseong Shin ⋅ Jihwan Eom ⋅ Jaeyeol Jeon ⋅ Sunghun Kang ⋅ Woo Young Kang ⋅ Seungryong Kim ⋅ Byungseok Roh
Abstract
Long-video inference in large vision-language models is bottlenecked by the number of visual tokens entering the language model, motivating compression under limited compute and memory budgets. Many methods prune internal activations or compress input features using decoder-agnostic heuristics. We analyze a frozen decoder with attention knockout and identify two layer-localized interactions: information formation among visual tokens and information selection by text tokens. Guided by these observations, we train a lightweight pre-LLM compressor with feature and relation distillation at the corresponding layers. The backbone remains frozen, and compression is query-agnostic and preserves FlashAttention compatibility. Across five video benchmarks, the compressor achieves the highest average accuracy among the compared compression methods at three retention ratios, with larger margins under aggressive compression. On the measured inference workload, it accelerates prefilling by up to $12.3\times$ and reduces peak memory by approximately $12\%$.
Chat is not available.
Successful Page Load