Beyond Fixed Memory: A Survey of Neural Memory Mechanisms and Test-Time Learning for Sequence Modeling
Abstract
The tension between computational efficiency and long-range sequence modeling remains one of the most consequential problems for large language models~(LLMs). Transformer architectures, while dominant due to their expressive attention mechanism, impose quadratic time and memory complexity with respect to sequence length, making them costly for contexts extending beyond thousands of tokens. Classical recurrent neural networks (RNNs) offer linear complexity but compress all historical context into a fixed-size hidden state, sacrificing recall as sequences grow. This survey reviews a rapidly emerging paradigm, neural memory mechanisms with test-time adaptation, that resolves this tension by treating the hidden state of a recurrent model as a learnable, dynamically updatable memory that continues to train during inference. We systematically examine recent landmark architectures, analyzing memory representation, computational complexity, training strategy, and empirical performance on long-context benchmarks, and further review the broader landscape of sub-quadratic alternatives to Transformers. We identify open challenges: training efficiency, optimal memory capacity, catastrophic forgetting, hardware alignment, theoretical foundations, benchmark maturity, and multi-modal generalization, and propose directions for future research. Beyond an architectural review, we determined how these mechanisms and challenges map onto the design problems already faced by long-term memory for conversational and personalized AI agents: catastrophic forgetting is the architectural mirror of what agent-memory systems call consolidation, and hardware/training efficiency determine whether test-time updates are viable at conversational latency. This survey is, to our knowledge, the first to unify the test-time memory literature with both the wider sub-quadratic sequence modeling field and the practical design space of long-term agent memory, offering researchers and practitioners a structured reference for selecting and extending memory-augmented architectures.