Video Token Compression via Formation and Selection Distillation in Large Vision Language Models
Abstract
Large Vision-Language Models face severe computational bottlenecks during long video inference due to an explosion of visual tokens. Existing token reduction methods either prune tokens dynamically within the LLM (intra-LLM), which breaks compatibility with FlashAttention and disrupts multi-turn conversations, or rely on LLM-agnostic heuristics before the LLM (pre-LLM), which can discard vital context. In this paper, we bridge this gap by analyzing the LLM's internal layers, discovering that visual processing consists of two distinct, layer-concentrated phases: information formation (vision-vision interaction) and information selection (vision-text interaction). Motivated by this, we propose a lightweight, trainable pre-LLM compressor that preemptively emulates these internal LLM behaviors while keeping the backbone frozen. We introduce two novel distillation objectives—vision feature alignment and vision-text relation alignment—to directly supervise the compressor at the layers where these internal interactions peak. Evaluated on video understanding benchmarks, our compressor achieves state-of-the-art accuracy-retention trade-offs against competitive baselines, showing the largest performance gains at aggressive token compression ratios.