VidHalluDoctor: Learning Video Differences to Mitigate Hallucinations in Vision-Language Models
Abstract
Large Vision-Language Models (LVLMs) show strong video understanding, yet hallucinations persist, with outputs that misalign with the video content. This stems from two primary causes. First, consecutive frames are often semantically similar, so the model can miss visual details. Second, LVLMs may rely on language priors and pay less attention to visual evidence. To mitigate video hallucinations, we propose VidHalluDoctor, a training framework that enhances fine-grained discrimination and reduces reliance on language bias. VidHalluDoctor builds video difference samples to improve visual detail recognition, counterfactual samples to reduce prior bias, and consistency-aware reinforcement learning with designed rewards that evaluate both reasoning quality and consistency between the reasoning and answer. To better evaluate video hallucinations, we introduce VidHalluEva, a comprehensive benchmark comprising 10,000 videos that covers three progressive hallucination types across diverse scenarios and multi-view settings. Extensive experiments demonstrate that VidHalluDoctor significantly reduces hallucinations while maintaining strong video understanding. Code and dataset will be released.