Understanding and Enhancing the Expressivity of Kimi Delta Attention
Julien Siems ⋅ Riccardo Grazzi ⋅ Jaisidh Singh ⋅ Korbinian Pöppel ⋅ Arber Zela ⋅ Timur Carstensen ⋅ Volkan Cevher ⋅ Antonio Orvieto ⋅ Aaron Klein
Abstract
Delta-rule linear attention layers, such as Gated DeltaNet (GDN) and Kimi Delta Attention (KDA), combine a diagonal decay gate with a generalized-Householder update. Since a diagonal matrix is itself a product of axis-aligned generalized Householder transformations, the full state transition is a product of such transformations. Prior work shows that widening the delta-rule learning rate enables parity and improves state-tracking. We show that widening GDN's scalar gate adds little, but widening KDA's diagonal gate to include $-1$, combined with the learning-rate extension allows a single KDA layer to go from parity to track every finite subgroup of $SO(3)$, including $S_4$ in three dimensions and $A_5$ in four. Moreover, three layers suffice to solve any group-word problem in finite precision, and, if the stability constraint $\beta \le 2$ is relaxed, to realize any weighted finite automaton in polynomial precision, one layer fewer than GDN. One-layer experiments on state-tracking tasks $S_3$ and $S_4$ confirm that both range extensions are required together. In language modeling, a 340M-parameter model with signed KDA is comparable with GDN, KDA, and DeltaProduct baselines.
Chat is not available.
Successful Page Load