Q-Focus: Let the Question Guide What to See in Long Videos
Abstract
Video Large Language Models face severe computational and memory bottlenecks in long-video understanding. Existing visual token compression methods predominantly rely on static, query-agnostic redundancy elimination. Consequently, they lack the ability to dynamically preserve query-relevant details, failing to emulate the human "coarse-to-fine" cognitive process. To bridge this gap, we propose \textbf{Q-Focus}, a plug-and-play, two-stage visual token compression framework. Q-Focus first constructs a holistic semantic representation via a global compression stage. This is followed by a question-guided focusing stage driven by two novel components: the Adaptive Contrastive Module (ACM) and the Question Context Module (QCM). Specifically, ACM achieves adaptive query-visual alignment by treating the user queries as positive samples and dynamically sampling negative samples from the video's feature distribution. Complementarily, QCM leverages visual affinity matrices to diffuse semantic relevance from highly matched visual anchors to surrounding contextual regions, effectively capturing implicit yet critical details. Extensive experiments demonstrate that Q-Focus seamlessly integrates with diverse compression baselines. Remarkably, while retaining only 10\% of the original visual tokens, it achieves relative accuracy gains of 7.3\% and 6.6\% on LLaVA-OneVision and LLaVA-Video, respectively, and consistently boosts the performance across other state-of-the-art models such as Qwen3-VL-8B-Instruct.