Retrieval-Augmented LLM Agents: Learning to Learn from Experience
Abstract
While large language models (LLMs) have advanced the development of general-purpose agents, robust generalization to unseen tasks remains challenging. Current approaches rely on either fine-tuning or training-free memory-augmented generation using retrieved experience; yet both have limitations: fine-tuning often fails to extrapolate to new tasks, while experience retrieval often underperforms compared to supervised baselines. In this work, we combine these approaches and study how retrieval-augmented LLM agents can learn to use retrieved trajectories in-context. First, we establish a strong LoRA fine-tuning baseline that outperforms several state-of-the-art agent training pipelines. Second, we analyze key design choices for experience retrieval, including storage, querying, and trajectory selection. Finally, we integrate experience retrieval directly into the fine-tuning process. Our results demonstrate that this substantially improves generalization to unseen tasks. We further show that these gains often persist with sparse, mismatched, and suboptimal experience, and even when agents reuse their own failed attempts, without parameter updates. Overall, simple episodic retrieval emerges as a strong foundation for agent memory, and shows that agents can be trained to learn how to use experience, rather than merely retrieve it. Code and Models will be made available upon acceptance.