Selective Cross-Modal Alignment in Vision-Language Models: Why Position Grounding Succeeds but Physical Reasoning Fails
Abstract
Physical intelligence—the ability to perceive dynamics and anticipate what happens next— is central to general intelligence, yet little is known about whether vision–language models (VLMs) internally represent it. We probe Qwen2.5-VL-3B on 6,000 fully controlled synthetic physics clips and uncover a selective cross-modal alignment failure: position is linearly decodable from visual tokens (R2 = 0.97) and perfectly grounded in language (100% success), whereas velocity is encoded only as an implicit, entangled manifold (R2 = 0.42, non-linear probe) that language generation does not use (19% physical-reasoning accuracy; less than 5% causal effect when the velocity-encoding layers are patched). This is not a missing capability: prompt engineering raises physical reasoning by +62% and eliminates a systematic response bias, while leaving the visual representations unchanged (cosine similarity ≈ 1.0). Comparisons with DINOv2 and SigLIP show that language training created the velocity encoding (R2 : 0.04/0.07 → 0.42), yet the default readout ignores it. Crossmodal alignment in VLMs is therefore selective and prompt-dependent: a model can hold physical knowledge it does not use, and capability must be measured by probing and activation, not by behavior alone.