Rethinking the Effectiveness of Contrastive Decoding in Mitigating Hallucinations in MLLMs
Abstract
Multimodal large language models frequently describe objects that are not present in an image. Contrastive decoding has become a widely used inference-time remedy, and it is assumed to work by contrasting a normal forward pass against a hallucination-prone one so that hallucinated content is suppressed. However, whether the reported benchmark improvements actually arise from this mechanism has not been examined. Our key observation is that a correction which suppresses hallucination must depend on whether the model is hallucinating, and that this dependence can be measured directly inside the model. Following this idea, we analyse the internal computation of three contrastive decoding methods on both discriminative and generative tasks, and compare each against controls that remove the contrastive term while preserving its effect on the output. The correction turns out to be unrelated to hallucination at every layer, the gains in captioning come instead from a constraint that narrows sampling towards greedy decoding, and an unstructured perturbation of the same magnitude performs at least as well. Experiments across three models and standard hallucination benchmarks show that these improvements reflect a change in decoding behaviour rather than genuine hallucination mitigation.