Training-free Focus-Ambient Retention for Memory-Efficient Video Large Language Models
Haozheng Zeng ⋅ Qi Bi ⋅ Gui-Song Xia
Abstract
Video large language models are often dominated by visual tokens, which lengthen prefill, enlarge KV-cache residency, and raise peak memory. In spirit of the philosophy $\textit{focus more and memorize less}$, we introduce Focus--Ambient Retention (FAR), a training-free visual token retention method that routes visual evidence through two complementary streams. The Focus stream preserves task-critical objects, actions, text, and fine details, while the Ambient stream keeps scene context only when it differs from a temporal cache. FAR combines visual attention, local frequency variation, and bounded query relevance to score tokens, then assembles a fixed-budget context with diversity control before language-model prefill. Across five video understanding benchmarks under a variety of model sizes, FAR significantly reduces peak VRAM by up to $50.6\%$ under matched retained-token budgets while achieving comparable or even superior performance over the state-of-the-art, improving the quality-memory trade-off. Overall, FAR offers a position-aware alternative to frame- or block-level compression for efficient inference. Code will be available.
Chat is not available.
Successful Page Load