Positional Encoding and Prompt-Order Sensitivity in In-Context Learning
Zhijie Wang ⋅ Bo Jiang ⋅ Shuai Li
Abstract
Transformer models have demonstrated a remarkable ability to perform a wide range of tasks through in-context learning (ICL), where the model infers patterns from a small number of example prompts provided during inference. However, empirical studies have shown that the effectiveness of ICL can be significantly influenced by the order in which these prompts are presented. Despite its significance, this phenomenon has been largely unexplored from a theoretical perspective. In this paper, we theoretically investigate how positional encoding (PE) affects the ICL capabilities of linear attention transformer models, particularly in tasks where prompt order plays a crucial role. We examine two distinct cases: linear regression, which represents an order-invariant task, and dynamical systems, a classic time-series task that is inherently sensitive to the order of input prompts. Theoretically, we evaluated the change in the model output when two types of positional encoding (one-hot and RoPE) is incorporated and the prompt order is altered. In all cases, the leading dependence on the permutation size and context length is $k/N$, with constants determined by the positional encoding, model weights, input dimension, and task dependent parameters. These theoretical findings are experimentally validated.
Chat is not available.
Successful Page Load