RoPE Is Not a Proper Relative Position Embedding
Abstract
Rotary Position Embedding (RoPE) is one of the most widely used positional mechanisms in modern large language models. Because RoPE applies position-dependent rotations whose attention score can be rewritten in terms of relative offsets, subsequent work has sometimes classified it as a relative position embedding. In this position paper, we argue that RoPE is not a proper relative position embedding. Treating RoPE as a relative position embedding can create overly strong expectations about length extrapolation, context extension, and comparisons with other positional mechanisms. In practice, however, these expectations have not been borne out: RoPE-based long-context models typically rely on position interpolation, continued training, or fine-tuning. This is more consistent with an absolute-position-conditioned view of RoPE. We argue that this reframing, which is closer to the original formulation, offers a more precise conceptual foundation for analyzing RoPE, understanding long-context behavior, and comparing positional mechanisms in modern LLMs. To make this distinction precise, we formalize the distance-decay property that has been implicitly expected of relative position embeddings and show that RoPE does not satisfy it.