TIMA: Test-Time Internalization for Agentic Memory
Abstract
Agentic memory is critical for accumulating and reusing experience in LLM agents across long interaction horizons and related tasks. While latent-memory approaches offer a compact representation of such experience without incurring additional context overhead, existing methods are designed to encode transient in-context states or rely on pre-trained modules that remain fixed at test time to inject experiential representations. Consequently, an effective mechanism for continually refining reusable latent memory from execution experience is still missing in current agents.To this end, we propose TIMA, a framework for test-time internalization that enables agents to acquire and update latent memory during interaction. TIMA couples a memory-augmented agent loop with an online-updated Internalizer and an Internalizer Bank for retrieval and continual refinement across episodes. Through self-supervised updates derived from execution trajectories, TIMA progressively encodes task-relevant solving patterns into compact latent memory, enabling both within-episode adaptation and cross-episode transfer. Extensive experiments across eight benchmarks show that TIMA consistently outperforms strong baselines, surpassing A-Mem by up to 50.8% and latent-memory baselines such as MemGen and Titans by up to 4.47%--15.96%. Further analysis indicates that the learned latent memory remains stable under continual updates, generalizes well out of distribution, and exhibits interpretable clustering structure across Internalizers.