V-CAST: Video Curvature-Aware Spatio-Temporal Pruning for Efficient Video Large Language Models
Xinying Lin ⋅ Xuyang Liu ⋅ Yiyu Wang ⋅ Teng Ma ⋅ Jiasheng Li ⋅ Zichen Wen ⋅ Wenqi Ren
Abstract
Video large language models (VideoLLMs) show strong capability in video understanding, yet long-context inference is still dominated by massive redundant visual tokens in the prefill stage. We revisit token compression for VideoLLMs under a tight budget and identify a key bottleneck, namely insufficient \textit{spatio-temporal information coverage}. Existing methods often introduce discontinuous coverage through coarse per-frame allocation or scene segmentation, and token merging can further misalign spatio-temporal coordinates under MRoPE-style discrete $(t,h,w)$ bindings. To address these issues, we propose V-CAST (\textbf{V}ideo \textbf{C}urvature-\textbf{A}ware \textbf{S}patio-\textbf{T}emporal Pruning), a training-free, plug-and-play pruning policy for long-context video inference. V-CAST treats frame representations as a semantic trajectory and uses local curvature as a low-cost proxy for temporal turns, routing token budgets to transition regions without relying on decoder attention. It further adopts a dual-anchor spatial selection mechanism that preserves high-entropy visual evidence without attention intervention, while keeping retained tokens at their original coordinates to maintain positional alignment. Extensive experiments across multiple VideoLLMs of different architectures and scales demonstrate that V-CAST achieves \textbf{98.6\%} of the original performance, outperforms the second-best method by \textbf{+1.1\%} on average, and reduces peak memory and total latency to \textbf{86.7\%} and \textbf{86.4\%} of vanilla Qwen3-VL-8B-Instruct. \textit{Our code is provided in the supplementary material.}
Chat is not available.
Successful Page Load