Proteus: Incremental Memory Activation for Long-Context Sequence Modeling
Abstract
The quadratic cost of attention-based sequence models for long contexts has motivated a growing line of research on memory-based models that can compress context into a compact state. However, most existing memory models expose a static memory throughout the entire sequence. In practice, this approach is prone to memory being “polluted” by early inputs, and it can saturate and fail to incorporate later context. We study a new paradigm of incremental memory activation, where the effective capacity of memory is controlled and progressively expanded as the context grows. By imposing an early bottleneck and introducing fresh capacity over time, this leads to better compression of history and reduces interference. We instantiate this paradigm in Proteus, a straightforward mechanism that can be incorporated into a broad class of neural memory architectures at no additional cost. We apply Proteus to state-of-the-art models, including Hope-Attention, SWLA, Comba, and Titans, and observe consistent improvements across all of them. Overall, our results show that static memory is suboptimal and that incremental memory activation is a promising direction for long-context management.