Agent Online Learning Beyond Memory
Abstract
Large language models are typically trained offline and remain fixed at deployment, even though users' preferences and environments can change over time. In-context learning with memory can adapt immediately from recent feedback, but its gains are often limited, while reinforcement learning can achieve stronger long-term adaptation but learns through costly trial and error. This creates a fundamental mismatch with deployment-time learning: unlike batch training, where only the final model matters, every mistake made during adaptation directly affects the user. We introduce Fast and Slow Reinforcement Learning (FSRL), which combines fast adaptation through memory with slower reinforcement learning updates that consolidate successful experience into the model weights. We evaluate FSRL across diverse settings spanning interactive decision making, tool use, factual answering, and personalized generation, under both stationary streams and distribution shifts where tasks or user preferences change over time. Across these settings, FSRL accumulates fewer errors while adapting than either memory or reinforcement learning alone, showing that combining fast adaptation with slower consolidation is critical for effective learning during deployment.