GEM-MM: Reducing VLM Hallucination with Entropy-Guided Preference Alignment
Yuanshan Chua ⋅ Xuejiao Zhao
Abstract
Vision-language models (VLMs) are routinely aligned with offline pairwise objectives inherited from text-only RLHF, but multimodal preference data is scarce and every response token is treated as equally informative. Under a fixed, small preference budget this makes the choice of pair construction nearly irrelevant: standard DPO-style multimodal objectives, image-conditional preference optimization, and self-improvement pair mining all land within 1.1 points of one another on held-out MM-RLHF preference metrics and within 0.5 points on generative hallucination rate, and the supervised control that initializes them falls inside the same band. We call this regime entropy-blind preference optimization: the learning signal never looks at where the policy is uncertain, even though in chain-of-thought VLM responses uncertainty concentrates on a small set of fork tokens at which the model commits to a visual claim. We introduce GEM-MM, which extends entropy-guided preference modeling to VLMs: it samples on-policy candidates, rewards high-entropy reasoning forks while penalizing uncertain final answers, and updates the full model with a group-normalized policy-gradient step. On a locked 3000-prompt MM-RLHF budget with Qwen3-VL-4B-Instruct, GEM-MM attains the best entropy-depth preference score (61.7% vs. 57.3–58.4% for the preference-tuned baselines) and the best objective-agnostic implicit-reward preference accuracy (41.4%), and it reduces AMBER generative hallucination from 23.4–24.0 to 20.6 and CHAIR from 5.0–5.1 to 4.6 at matched caption coverage. It is also the best system on all three HallusionBench granularities, while a conservative operating point of the same objective raises the near-chosen rate to 62.4%. An off-the-shelf multimodal reward model that never saw our objective ranks GEM-MM first among the systems compared, and two such reward models independently confirm our two main ablations at $p < 0.001$. We report per-metric confidence intervals and paired significance tests, state explicitly which of our metrics share functional form with the training reward, and document a collapse mode that appears when continue-training is pushed past the fair budget; code, configurations, and the evaluation harness are released.
Chat is not available.
Successful Page Load