Reading in Bands: Causal Maps of Long-Context Consumption in Vision-Language Models
Abstract
Long-context foundation models increasingly read compressed or cached context, but capacity and token counts do not reveal where a decoder can use it. We study this demand side with rendered-page compression, which roughly halves decoder tokens while exposing a retrieval–reasoning gap. We introduce two training-free attention interventions on frozen vision-language models. A read gate blocks question-to-context attention at selected decoder layers; a supply gate blocks direct off-diagonal attention among visual-token positions while leaving question access open. All reported comparisons use one explicit-mask attention path: repeated open controls are bit-identical, and full read blocks are exactly invariant to replacing the gated visual embeddings with noise. We apply the read gate to Qwen2.5-VL-7B/3B and GLM-4.1V-9B, and the supply gate only to Qwen2.5-VL-7B. Read-gate maps reveal model-specific depth structure with sharp lower edges: Qwen-7B retains 97–98% of retrieval ceiling using only its top half, whereas GLM retains 9% there and reads through an interior band. On Qwen-7B, blocking direct visual-token interactions leaves retrieval intact (59/59 full-block item retention) but reduces association retention to 13/47. Read and supply blocks also produce contrasting errors: the former removes context names, while the latter usually selects the wrong name from the same context. Restricted read bands are associated with much longer GLM traces, but extra decoding does not restore failed access. The resulting maps provide a causal screen for layer-selective context-cache designs: they reject harmful layer sets before systems work and nominate model- and task-specific sets for physical-eviction tests. The result is a practical audit of long-context use rather than another capacity benchmark.