TeMPO: Frame-Causal Token Compression for Efficient Video Large Language Model
Abstract
Video large multimodal models incur substantial inference cost because visual tokens accumulate across sampled frames. Training-free token compression can reduce this cost, but it often degrades video reasoning performance when later frames must contribute new evidence over time. We identify a common failure mode behind this degradation: existing compressors do not track the committed support, namely the token support already retained and delivered to the language model by earlier frames. As a result, selectors with different scoring rules repeatedly retain overlapping supports across neighboring frames, a phenomenon we call Selector Collapse. To address this issue, we propose \tempo{}, a frame-causal token compression framework that combines committed-support memory, residual-energy greedy selection, and parameter-free neighborhood fusion. Across four video understanding benchmarks and multiple backbones, \tempo{} consistently improves the accuracy-efficiency trade-off of training-free compression; at 10\% retention, it preserves 99.8\% of vanilla performance on LLaVA-OneVision and improves fixed-budget frame scaling on Qwen2.5-VL to 109.3\% relative accuracy, without retraining, backbone modification, or additional token budget.