StreamMind: Dynamic Streaming Cognition for Online Video Understanding
Abstract
Online video large language models (VLLMs) have demonstrated remarkable capabilities in real-time streaming video understanding and proactive human-AI interaction. However, existing methods often follow a static perception paradigm, processing video streams with rigid temporal units while overlooking the inherent temporal non-uniformity of streaming videos. This oversight results in flickering responses and fragmented context. In this paper, We propose \textbf{StreamMind}, a dynamic streaming cognition framework that organizes both online reasoning and memory around temporally coherent Group-of-Pictures (GoP) segments. Specifically, StreamMind introduces two key components. First, StreamMind introduces a \textbf{Dynamic Thinking (DynThink)} module, which accumulates evidence within each GoP and triggers reasoning only at GoP boundaries, thereby reducing premature decisions and improving response stability. Second, to support long-horizon streaming understanding, we propose a \textbf{Brain-Eye Synergy Memory (BESM)} module, which preserves key-frame visual tokens as high-fidelity anchors while compressing the remaining frames into compact thinking tokens. This design maintains fine-grained visual evidence together with reasoning continuity, while reducing redundant storage. Extensive experiments on streaming and offline video benchmarks demonstrate that StreamMind achieves state-of-the-art online video understanding performance, reaching 63.4\% on OVO-Bench and 80.2\% on StreamingBench, while maintaining strong generalization to offline long-form video reasoning.