Agent Online Learning Beyond Memory
Abstract
During deployment, large language model (LLM) agents continuously generate new experience that can help them adapt to user preferences, task domains, and changing environments. Unlike conventional training, deployment-time learning is an online learning problem: intermediate checkpoints serve the users at every step, so a learner should be evaluated by cumulative performance over every interaction. Memory adapts quickly by reusing experience but often plateaus, whereas reinforcement learning (RL) can achieve stronger long-run performance but requires many samples to improve. We introduce \emph{Fast and Slow Online Learning (FSOL)}, which combines fast contextual adaptation through memory with slower parametric learning through RL. Across writing-style adaptation, search-augmented QA, WebShop, and ALFWorld, FSOL matches the rapid early gains of memory while approaching the higher performance ceiling of RL, yielding substantially better cumulative performance, especially under distribution shift. FSOL can retain previously learned capabilities while acquiring new ones under sequential training.