SMat-Attention: Structured Long-Context Sequence Modeling
Abdullah Ateyeh ⋅ Archer Wang ⋅ Emile Anand ⋅ Marin Soljacic
Abstract
Long-context sequence models face a fundamental tradeoff: softmax attention uses flexible token-level interactions at quadratic cost, whereas linear attention obtains linear-time training and constant-time decoding by compressing history into a fixed-size state. In this work, we ask whether we can connect these regimes through a tunable notion of structure. Therefore, we introduce SMat-Attention via a family of causal masks with structured long-range routing whose row supports have VC dimension $d$. In our construction, $d=1$ recovers the standard causal mask, and increasing $d$ permits richer subset-routing patterns. Moreover, we give chunkwise forward and backward algorithms to enable hardware-efficiency. For sequences of length $T$, we show that this mechanism takes $O(T^{2-3/d}+T)$ work, yielding $O(T)$-attention for $d\leq 3$ and $O(T^{2-3/d})$-attention for $d>3$, despite the mask being dense. For fixed-horizon streaming, decoding after the distant prefix takes time independent of $T$ per token using $O(T^{1-1/d})$ cached states. SMat-Attention therefore makes VC dimension an explicit knob governing access-pattern complexity, prefill cost, and decoding memory. Applying our SMat framework to Mamba-2 and Gated DeltaNet yields improvements on tasks such as subset-routing, multi-query associative recall, multi-item retrieval, selective copying, and performs competitively in small language modeling settings.
Chat is not available.
Successful Page Load