Reinforced Evidence-Aware Long Video Understanding
Abstract
Despite strong benchmark performance, multimodal large language models (MLLMs) remain prone to hallucinations, as they are optimized for linguistic plausibility rather than faithful answering based on observation. This issue is particularly severe in long videos, where critical cues are temporally sparse and often missed by uniform frame sampling. Yet existing methods still either operate on fixed pre-sampled inputs, implicitly assuming that all necessary evidence has already been captured, or augment reasoning with retrieval without requiring answers to be grounded in visual evidence. We propose Reinforced Evidence-Aware Learning , a framework that reformulates long-video question answering as an iterative process of evidence gathering and verification, enabling faithful answering with explicit evidence support, thereby reducing hallucinations. At each reasoning step, the model assesses whether the current observation is sufficient to answer the question, and either searches for missing visual evidence from the video or produces a final answer with explicitly cited evidence. To internalize this reasoning policy, we design an evidence-aware reward for reinforcement learning post-training that jointly requires answer correctness and evidence faithfulness. The latter is evaluated by a pretrained cross-modal consistency verifier that matches the generated evidence descriptions against the visual content of the source frames, eliminating the need for manual annotations. Built on Qwen2.5-VL-7B, REAL achieves state-of-the-art results on four long-video benchmarks (Video-MME, MLVU, LongVideoBench, and EgoSchema) and substantially suppresses hallucinations on the dedicated video-hallucination benchmarks VideoHallucer and ELV-Halluc, with only 96 input frames surpassing the 768-frame base model in both accuracy and hallucination resistance.