Discovering Semantic and Structural Heads in Graph Transformers
Abstract
Graph Transformers (GTs) have achieved strong empirical results in graph machine learning, yet how they compose semantic and structural information internally remains poorly understood. For Large Language Models (LLMs), attention heads have been observed to specialise into positional and semantic roles. We propose an architecture-agnostic donor-swap methodology to test whether this dichotomy transfers to GTs. Our approach intervenes at the graph-input, scoring each head by how its output responds when semantic or structural information is swapped between nodes. We identify specialised selector and router motifs in all three analysed molecular checkpoints, spanning both tested GT architectures. However, the strict semantic-positional division reported for LLM attention does not transfer: semantic and structural responses are spatially co-organised, and heads combine the two signals, using structure to delimit where to look and semantics to determine what to select.