What Should Titans Forget? Disentangling Magnitude, Distribution, and Assignment in Test-Time Memory
Beloslava Malakova
Abstract
Test-time neural memory must continually balance retaining useful information against forgetting previous updates. Titans addresses this through learned, data-dependent decay, yet it remains unclear which properties of this mechanism account for its effectiveness. We study learned forgetting through controlled test-time interventions on a Titans Memory-as-Context model, decomposing forgetting into its magnitude, distribution, and content-dependent assignment. Removing decay severely degrades next-token prediction, while hand-designed policies based on surprise, novelty, recency, and their combination also substantially underperform native decay. Matching their mean decay recovers considerable performance, indicating that the overall forgetting regime is an important confounder when comparing retention strategies. More strikingly, shuffling the exact native decay values while preserving their per-layer empirical distributions increases loss by only $0.57\%$. Nevertheless, native assignment outperforms shuffled decay on $93.2\%$ of $1{,}954$ paired evaluation sequences. These results suggest that appropriate global decay dynamics account for much of the effectiveness of learned forgetting in our setting, while content-dependent assignment provides a smaller but highly consistent additional benefit.
Chat is not available.
Successful Page Load