Position Without Positional Embeddings: A Directional Mechanism in NoPE Transformers
Abstract
Autoregressive transformers with no positional embeddings (NoPE) can recover absolute position. A common explanation is that causal attention creates a position-dependent variance signal, but this account is incomplete because LayerNorm removes per-token scale, so a viable mechanism must encode position in direction, not magnitude. We describe the most compact mechanism with two-layer NoPE transformers that accurately extract position. In Layer1, prefix averaging creates a shared beginning of sequence (BOS)-direction component. In Layer2, the attention's output-value matrix maps the \bos{} and \nbos{} contributions into two distinct directions, and a position-dependent \bos{} attention weight acts as a mixing coefficient that interpolates between them. The resulting directional trajectory is linearly decodable for position. Interestingly, we show two transformer variants that converge to the same theoretical mechanism and instantiate it when trained to predict position.