Object Hallucination Mitigation in Large Vision-Language Models via Self-Vision Dual Masking and Uncertainty-Triggered Assembly
Abstract
Large vision-language models (LVLMs) achieve strong multimodal reasoning performance but remain prone to object hallucination, where generated objects or existence judgments are not faithfully grounded in the image. Existing mitigation methods either require additional training or rely on external visual experts, which limits plug-and-play usability and increases inference overhead. In this paper, we propose SIGMA, a training-free framework for hallucination mitigation based on Self-vision Information-theoretic Gating and dual-Mask Assembly. SIGMA operates directly on the LVLM's own vision encoder, constructing complementary support and complement masks from image-driven concept relevance, question-aware entity cues, and CLS--patch shortcut suppression. These masks form global, support-masked, and complement-masked visual streams for query-conditioned evidence separation. During decoding, SIGMA activates correction only under high uncertainty, using full-vocabulary and candidate-answer entropy to perform support-first refinement and complement-on-demand suppression. The resulting streams are combined through logit-level assembly with lightweight positive calibration and a plausibility constraint. Experiments across multiple LVLM backbones and hallucination-oriented benchmarks show that SIGMA effectively reduces object hallucination while maintaining competitive perception performance and a selective low-overhead inference path.