When Does a VLM Believe the Caption Over Its Eyes?
Abstract
This initial empirical study examines how a frozen vision-language model, Qwen2-VL-2B-Instruct, answers a left/right question about synthetic two-object images when a caption that agrees with or contradicts the image is placed in the prompt. Across matched caption conditions, answer behavior tracked agreement between the caption’s relation word and the correct answer more closely than whether the caption was true: a truthful swapped-subject caption and a false same-subject caption, both containing the opposite relation word, produced identical predictions, with the first option of the answer instruction selected on all 48 items. Random, greedy and MAP-Elites search were compared within a 51,840-template space using 30 candidate evaluations per run and two seeds per method; greedy search found templates that flipped every held-out clean-correct item, so coverage saturated at the tested budget. Activation patching with self-patch, repeat and unrelated-source controls showed that neutral-source image-token patches did not restore correct answers in any of 12 cases, whereas post-image text patches recovered answers from intermediate layers, and object color remained decodable from image tokens. Both greedy runs achieved 100% coverage of the evaluated held-out items, yet their five selected templates produced one versus five distinct binary failure vectors. Coverage alone therefore did not distinguish these sets of attacks.