Two Layers of Attention Stability: A Koopman-Operator Analysis of Linear and Softmax Attention
zhenglong chen
Abstract
We develop a Koopman-operator analysis of self-attention that resolves a long-standing puzzle: why does linear attention diverge on long sequences while softmax attention remains stable, even though both share the same next-step prediction objective? We show that attention stability is intrinsically a *two-layer* problem. The *architectural layer* concerns the spectrum of the $n \times n$ token-mixing matrix and is data-independent; the *dynamical layer* concerns the spectrum of the $d \times d$ dual operator induced by training and should be data-determined. We prove three concrete results within this framework: (i) under a shared encoder, the dual operator $C^* = W_Q W_K^\top G W_V$ of linear attention coincides with the EDMD Koopman estimator $K_{\text{EDMD}} = G^{-1}A$ at the global optimum, and its spectral radius is unconstrained; (ii) causal softmax attention enjoys an architectural guarantee $\rho(A^{\text{soft}}) \le 1$ via row-stochasticity, while its value projection $W_V$ leaves the dynamical layer entirely unconstrained; (iii) LayerNorm makes the empirical Gram matrix $G_{\text{LN}}$ rank-deficient with kernel $\text{span}(\mathbf{1})$, producing a $d$-parameter family of loss-equivalent but spectrally distinct EDMD solutions, while RMSNorm preserves invertibility. The framework yields an immediate design principle—constrain the architectural layer, free the dynamical layer—which we instantiate as two stabilized linear-attention variants (SN-LA, RN-LA). Numerical experiments reveal a *hierarchy* of architectural stability: vanilla linear attention has neither bound and diverges immediately under autoregressive rollout; our stabilized variants enforce only the single-pass spectral bound $\rho(\tilde{A}) \le 1$ and delay divergence; softmax attains the strictly stronger pointwise contraction $\Vert A v\Vert_\infty \le \Vert v\Vert_\infty$ via row-stochasticity and remains bounded. On real time-series benchmarks (ETT, Weather), our methods close $\ge 80\%$ of the vanilla-to-softmax gap and surpass softmax on $3/5$ datasets at long horizons, while preserving the dynamical-layer freedom that constraints on the dual operator forfeit.
Chat is not available.
Successful Page Load