Transfer Entropy as a Measure of Information Flow in VLMs and LLMs
Jessica E Liang ⋅ Jianbo Shi
Abstract
Understanding how information propagates within vision-language models (VLMs) and large language models (LLMs) is central to interpretability and efficient deployment. Attention weights indicate where probability mass is allocated, but do not directly measure directional influence, redundancy, or functional importance. We introduce transfer entropy (TE) as an information-theoretic measure of directed information flow in Transformers, and develop tractable TE estimators for high-dimensional hidden states at the granularity of layers, tokens, and attention heads. We first analyze multimodal models, beginning with LLaVA-1.5-7B and then CLIP ViT-B/16. In LLaVA-1.5-7B, vision$\rightarrow$text TE increases with depth, suggesting that multimodal fusion is localized in the upper decoder layers. In CLIP, TE reveals late-layer redundancy in the vision tower and broader elevated TE in the text tower. We then analyze unimodal models, including RoBERTa, T5, Llama-3.2-3B, and Qwen-2.5-7B. Across these LLMs, TE reveals consistent depth-dependent structure: RoBERTa exhibits a mid-layer peak, decoder-only LLMs concentrate task-dependent computation in broad mid-depth bands, and T5 shows distinct encoder and decoder dynamics. Beyond layerwise analysis, TE also exposes token- and head-level communication roles, distinguishing sinks, broadcasters, and benign components. These signals support TE-guided layer, token, and attention-head pruning across VLMs and LLMs, consistently outperforming attention- or saliency-based heuristics under matched budgets. Overall, TE provides a unified directed-information lens for diagnosing, interpreting, and compressing Transformer-based models.
Chat is not available.
Successful Page Load