Linear Attention as State Editing: Rethinking Error and Write Addresses
Abstract
Linear attention is derived by removing the softmax in the attention operation. It is seen as a recurrent update of a fixed-size state. This motivates the usual key--value view, which encourages fixed roles for learned projections and hides other possible formulations. We describe these recurrences instead as state editing: passive decay followed by one addressed rank-one edit. The edit has an error address, a write address, a target, and a strength. Our framework places DeltaNet, GDN, KDA, GDN-2, and Q-Delta in one design space. It also motivates \edm, an extension of GDN-2 that learns separate key--query mixtures for the error and write addresses. In matched 340M-parameter experiments, \edm{} improves language-model validation loss over GDN-2 and a competitive \qdstar{} baseline in both tested seeds. Retrieval improves in some comparisons but varies substantially across seeds.