Fast Weights Recurrent Transformer
Ivan Anokhin ⋅ Daria Yasafova ⋅ Tianyu Zhao ⋅ Irina Rish ⋅ Sebastian Risi
Abstract
The key--value (KV) cache of softmax attention can be viewed as a nonlinear fast-weight network, yet standard sliding-window caching updates it through first-in-first-out (FIFO) replacement.By dropping the requirement for sequence-parallel training, we build on recurrent Transformers and replace FIFO updates with sparse, content-dependent gated refinement of a fixed-size KV memory.On FineWeb, learned refinement reduces validation loss significantly relative to FIFO-based recurrent Transformers at tested memory sizes up to $128$ entries. On associative-recall and needle-in-a-haystack tasks, it learns faster and supports successful retrieval in settings where FIFO-based baselines fail, respectively.
Chat is not available.
Successful Page Load