Steering Object Faithfulness in VLMs with Vision-to-Language Transfer Vectors
Abstract
In this extended abstract, we study a derived class of \textbf{neural network artifacts}: the language-side displacement that each individual vision feature induces in a vision-language model (VLM). For every vision feature, we ablate it on its top-activating images and record the average difference between the intact and ablated representations in an early LLM residual stream. Collecting these displacements yields a bank of one vector per vision-feature dimension---an ablation-derived trace of how vision-side units act on language-side states for each model, distinct from activation magnitude alone. We then use this artifact directly as a source of inference-time intervention directions: our language-side compensation method adds an image-conditioned weighted sum of the stored vectors to that residual stream, steering the trade-off between object faithfulness and coverage in generation. Across four VLMs evaluated on MS COCO captioning, compensation yields contiguous ranges of tested strengths with higher target-object recall than a five-prompt verbosity-control frontier at matched hallucination rate (CHAIR-I). Directly amplifying vision-feature activations produces mixed results, and whether all features or the lowest-effect 10% give the clearer favorable range depends on the model, as do the shapes of the measured effect-magnitude distributions. These findings suggest that ablation-derived effect traces are worth computing, storing, and reusing as a \textbf{model artifact}, and that they supply useful image-conditioned intervention directions for object faithfulness.