Tree Rotary Positional Encoding for Extreme Length Extrapolation from Scratch
Abstract
Large language models rely on positional representations for long context modeling, yet are typically trained on short sequences and evaluated at much longer lengths. Bridging this gap requires positional encodings remaining distinguishable at large distances while generalizing beyond the training length. In this paper, we identify a structural limitation of existing rotary-based positional encodings: all positional dimensions are controlled by a single global index, which couples training-time positional variation with long-range extrapolation. To address this, we propose Tree Rotary Positional Encoding (TPE), a structured multi-scale positional encoding designed for training from scratch that represents absolute positions through com positions of different components, avoiding single-index extrapolation. We further devise an Offset-based Positional Training (OPT) strategy to improve learning interactions across TPE components, enabling better generalization to longer con texts. We theoretically analyze that its multi-scale decomposition enables effective combinatorial position coverage. Experimental results show strong gains under extreme length extrapolation: with a 1.2B model trained from scratch on length 512, our TPE achieves a 64× extrapolation to 32K tokens, reducing perplexity from 4532.21 (RoPE) to 149.96, with comparable inference-time efficiency.