When Dynamic Scaling Fails: A Scale–State Feedback Governs Recurrent-State Quantization in Linear-Attention and SSM LLMs
Abstract
Linear-attention and state-space LLMs carry long context in a fixed-size recurrent state, and quantizing that state is the natural way to cut its cost during long decode. Unlike a weight tensor, the state is rewritten at every step, so a quantizer applied at each step sits inside the recurrence and its errors compound. Recurrent-state quantization therefore fails along an axis that is almost never reported. A bit-width certified as safe at short exposure need not stay safe as decode proceeds, and INT4 that leaves synthetic recall perfectly intact at the start destroys it a thousand steps later. The mechanism is a feedback between the quantizer scale and the state it is derived from. Yoking that scale to a full-precision reference removes the resulting norm explosion entirely, yet recovers only 1.8 of the 5.4 nats lost, so numerical instability and functional degradation are separable failures. The remainder is a resolution problem that no cleverer bit allocation solves at matched rate: plain persistent group-wise INT4 dominates, and under article-disjoint splits a learned per-layer allocation collapses onto the trivial rule of protecting the first few layers. Run out to 32K, that winning policy is worth no more than discarding the state. What survives is a design principle. Contain the scale rather than locate outliers, and report the decode horizon at which a bit-width was certified.