RAM-Net: Linear-Time Sequence Modeling with Sparsely Addressable State
Kaicheng Xiao ⋅ Haotian Li ⋅ Liran Dong ⋅ Guoliang Xing
Abstract
Linear attention offers an efficient alternative to full attention with a fixed-size recurrent state. However, this state is shared by all tokens, so information from distinct tokens becomes superposed within it and produces inter-token interference that degrades long-range fine-grained recall. To address this issue, we propose RAM-Net, which replaces dense access to a shared state with sparse address-based access. RAM-Net organizes the recurrent state as a fixed-size array of independent slots and uses an Address Decoder that maps each key or query into a sparse address, selecting a small subset of slots to write to or read from at each step. This design directs tokens with non-overlapping addresses to disjoint slots, suppressing inter-token interference, while keeping per-step overhead dependent only on the number of accessed slots rather than the total state size. Empirically, RAM-Net shows a clear advantage over strong linear baselines on fine-grained long-range retrieval and remains competitive on standard language modeling and commonsense reasoning, while accessing far fewer state elements per step than these baselines (e.g., 32$\times$ fewer than Mamba2).
Chat is not available.
Successful Page Load