Titan-Llama: Retrofitting Pre-trained LLMs with External Neural Memory Modules
Sarah Pan ⋅ Ali Behrouz
Abstract
Efficient long-context architectures often improve inference at the cost of requiring models to be trained from scratch. We introduce Titan-Llama, a post-training method that retrofits pretrained transformers with hierarchical memory, preserving full attention over recent tokens while compressing historical context into a fixed-size neural memory module. On Llama-3.2-1B-Instruct, Titan-Llama recovers substantial performance lost to attention segmentation, exceeds full attention on RULER multi-key needle-in-a-haystack tasks, and achieves over 80$\times$ faster prefill and 20$\times$ lower steady-state decode memory usage at long context lengths.
Chat is not available.
Successful Page Load