When Semantics Survive but Symbols Go Stale: Response Conventions in Vision–Language Models
Abstract
Foundation models must not only recover a correct proposition; they must express it through the response convention currently licensed by the dialogue. We test whether vision–language models (VLMs) can preserve stable visual semantics while replacing a linguistic convention that maps those semantics to answer sym bols. Across four VLM checkpoints, we conduct 3,584 condition-level behavioral evaluations under a controlled same-image repeated-measures protocol: a true caption and a minimal foil exchange option labels between turns, while the image, captions, and task remain fixed. Under a correct full-history turn, Qwen3.5-4B reaches 100% accuracy when the true caption moves from A to B, but only 31.25% in the reverse direction, despite 90.63% no-history accuracy. Symbol-free probes remain 95.83% accurate, showing that propositional content survives while symbol emission fails. Clean-to-corrupt activation patching fully rescues selected failures at the final layer; removing prior option content raises the failing direction to 96.88% without harming the reverse direction. A newer Qwen3.8-27B solves the exact swap, yet a 2×2 transcript intervention shows that the preceding requested format changes how an identical experimenter-injected correct sentence affects the next answer (B→A accuracy interaction +25 percentage points, bootstrap 95% CI [12.5,40.63]). Exploratory replications yield positive B →A margin interac tions in Gemma-3-4B (+2.920, CI [1.026,4.831]) and Llama-3.2-11B (+1.175, CI [0.775,1.596]), while absolute cell accuracies reveal model-specific floor and ceiling effects. The results identify a semantic–symbol dissociation and show that response form and its relation to a prior request are substantive parts of multimodal computation rather than interchangeable output wrappers.