Reading in Bands: Training-Free Causal Maps of Where Vision-Language Models Consume Compressed Context
Lakshya Narula
Abstract
Rendering text as page images lets a vision-language model carry context in roughly half the tokens, making pixels an alternative reading medium to subword tokens. The medium has a known asymmetry: literal retrieval survives it, but reference resolution, binding an attribute to the entity that owns it, does not. We ask where in the decoder rendered context is consumed, and what the context must do to itself first. Two training-free attention gates answer both causally. A read gate blocks the question from attending to the context at chosen layers; a supply gate blocks context tokens from attending to one another. Both are verified at the logit level: with nothing blocked, they are bit-inert, and with the read gate fully closed, replacing the whole context with noise changes the logits by exactly zero. Across Qwen2.5-VL-7B, Qwen2.5-VL-3B, and GLM-4.1V-9B, the gates reveal three shared regularities: each model reads through a band of layers with a two-layer onset; access from outside the band is absent rather than degraded; and the band begins earlier for tasks the model finds hard. Band location is model-specific: Qwen reads only from mid-depth upward, whereas GLM reads through an interior band and retains 72% of its retrieval ceiling when its final quarter is blocked, whereas Qwen scores zero. On Qwen2.5-VL-7B, the supply gate yields a double dissociation: severing all intra-context attention leaves retrieval untouched (58–60/60) but collapses two-hop association (47 to 13–24, $p<10^{-5}$), and the two failures differ in kind: a blocked reader answers from the question alone, whereas a blocked span names the wrong character. Consumption maps are cheap and model-specific: compressed context should be served where a model reads.
Chat is not available.
Successful Page Load