When the Fix Becomes the Flaw: Prompt-Prefix Conditioning Induces Visual Neglect in Multimodal Video Captioning
Abstract
We investigate a fundamental faithfulness failure in prompt-guided Vision-Language Models (VLMs): textual conditioning can cause the decoder to ignore the visual modality entirely, generating plausible-sounding but visually ungrounded captions. We term this Visual Neglect (Lazy Decoder effect). Our study proceeds through three architectural phases on 52,962 human-annotated video-caption pairs. We first show that a naive implicit baseline fails at fine-grained facet recognition (Emotion: 16.7%, Style: 8.2%). Attempting to fix this with joint multi-task auxiliary heads triggers Gradient Competition, where classification gradients corrupt the shared ViT encoder, causing decoder collapse (Emotion drops further to 11.3%). We introduce Stop-Gradient and Prompt-Prefix Conditioning to resolve gradient interference, recovering facet accuracy by up to 13.8 percentage points. However, a controlled four-condition counterfactual ablation reveals a critical trade-off: the model generates coherent captions from a zeroed (black-frame) visual input, and follows a corrupted prompt over real visual evidence in 100% of tested samples. This constitutes a form of modality hallucination that standard NLG metrics cannot detect.