Not All Attention Is Created Equal: Robustness of In-Context Learning Across Attention Mechanisms
Ken Zheng ⋅ Joshua Lu ⋅ Lenci Ni ⋅ Tiger Zhang ⋅ William Li
Abstract
We study the robustness of in-context learning (ICL) for linear regression under distribution shift. We evaluate seven attention mechanisms---including softmax, multi-query attention (MQA), local-global attention, and efficient non-softmax variants (ReLA, FAVOR, ReBased)---across three parameterized shift types: task-level, prompt-level, and query-level. Non-local mechanisms (softmax, MQA, FAVOR, ReLA, ReBased) generally maintain lower error under task-level and most query-level shifts. Under prompt-level shifts, however, performance becomes shift-dependent: ReLA and ReBased can outperform softmax and MQA when inputs are restricted to very low-dimensional subspaces, and local-global attention with a small window ($\ell_{\text{window}}=5$) is also competitive in this regime. Overall, no single attention mechanism is uniformly best; robustness depends on both the type and magnitude of the shift.
Chat is not available.
Successful Page Load