Rethinking Long-Video Efficiency: A Joint Allocation Perspective on Frames, Pixels, and Front-End Latency
Abstract
Efficient long-video understanding with vision-language models (VLMs) has largely been treated as informative token selection at fixed native resolution: which frames or which visual tokens to retain under a token budget. We argue this framing leaves two axes unused --- per-frame resolution itself can be traded for denser temporal coverage, and front-end decoding latency scales with the candidate pool, not the token budget. Through an empirical study across multiple VLMs and long-video benchmarks, we distill three lessons: (i) dense low-resolution sampling outperforms sparse native-resolution sampling at matched token budgets; (ii) some tasks are resolution-sensitive and benefit from high-resolution frames; and (iii) front-end decoding dominates wall time on hour-scale clips. These lessons motivate LoHi, a training-free, single-pass framework that pairs a dense low-resolution video stream with a sparse set of high-resolution image streams, processed through the VLM's native video and image pathways. Two plug-and-play selectors choose Hi-I frames at near-zero or low overhead: LoHi-Anchor uses codec-level I-frame metadata, and LoHi-SemDiv uses a query-relevance and visual-diversity DPP over CLIP features. On three long-video benchmarks, LoHi improves over the vanilla native-resolution baseline by +10.6% on average at matched token budget and over the strongest prior efficiency methods by +5.2%, while reducing front-end decoding latency by up to 7x on hour-scale clips.