Impact of Active Linguistic Streams on Visual Representational Geometry in VLMs
Abstract
Vision-language models (VLMs) that encode image patches and text tokens suffer from representational crowding: as the number of objects in a scene grows, their token representations become increasingly similar. An approach to mitigating crowding could involve the inclusion of language tokens to help detangle the set of visual object representations. We present three experiments that investigate this approach. First, we replicate and extend the static crowding effect in a standalone vision encoder, SigLIP-SO400M, showing a descriptive monotonic ordering of object-token similarity by image object-density across mid-to-late layers. Second, using Visual Genome-derived spatial relation queries, we show that this crowding is weakly but consistently associated with downstream answer accuracy across two encoders and two complementary statistical tests, and that a VLM's vision tower is already more crowded than its vision-only counterpart even absent of any text input. Third, we causally manipulate the content and length of the accompanying text prompt and measure its effect on PaliGemma's object-level visual-token geometry. We find that while semantic prompt relevance did not cause consistent, significant effects, prompt length significantly increased RDM distortion and Procrustes distance and reduced intrinsic dimensionality. Together, these results indicate that VLM visual-token geometry is not independent of the language context provided. Experimentally varying text prompt length alters multiple aspects of PaliGemma's internal visual-token representations; however, we find limited evidence for a global effect of the semantic relevance of the textual prompt.