Transformers with Slow-Growing Memory for Long Context
Abstract
Transformer architectures have been established as the de-facto backbone of recent advances in deep learning. However, their key component–i.e., attention mechanism–has a linearly growing memory that comes with a quadratic computational cost with respect to the sequence length, prohibiting Transformers from scaling to large context. More efficient alternatives, such as recurrent models use a fixed-size memory, making their computational cost linear with respect to sequence length. Despite their efficiency, they underperform Transformers in recall intensive tasks. In this paper, we build upon a new perspective on compressing the KV-cache of the Transformers into slow-growing blocks, where the model uses a set of latent parameters to compress its own KV cache into a sublinearly growing memory, allowing for arbitrary growth rate for KV-cache in Transformers. To support our design, we perform experimental evaluations on language modeling, common-sense reasoning, and long-context tasks, showing that our Transformers with slow-growing KV cache almost recovers the performance of global Transformers, while requiring significantly less compute.