No Reset, Emergent Behaviors: Evaluating and Training Agents on Sequences of Tasks
Romain Froger ⋅ Matteo Bettini ⋅ Theophile Laurent ⋅ Kavya Kopparapu ⋅ Grégoire Mialon ⋅ Amine Benhalloum ⋅ Djame Seddah ⋅ Thomas Scialom
Abstract
AI agents are typically evaluated one task at a time, yet deployed agents operate in persistent sessions where conversation history, memory, files, and prior experience carry across tasks. We study this mismatch by introducing two sequential benchmarks, adapted versions of GAIA2 and SWE-Bench Pro, in which agents solve sequences of tasks within a single persistent session. Persistent context does not consistently improve task success: Opus 4.6 and GPT-5.4 remain near their independent baselines, whereas DeepSeek v4 Pro loses 5.9 percentage points on average. Even when average performance remains stable, consistency ($\mathrm{pass}^K$) declines across rollouts. Persistence also substantially changes agent behavior: agents use up to $60\%$ fewer output tokens, repeat less exploratory work, and use Python up to $4\times$ more often to orchestrate and extend tools. These changes reveal a central trade-off: persistent sessions enable in-context adaptation and reuse, but also cause weaker models to reduce effort too aggressively or propagate harmful state across tasks. We introduce Sequential RL (Seq-RL), which trains on persistent multi-task rollouts using local and downstream rewards. Seq-RL substantially narrows the sequential gap on GAIA2 Search, reverses it on Execution, and transfers zero-shot to SWE-Bench Pro without software-engineering training data, outperforming Single-task RL under sequential evaluation. In short, our results identify reliable adaptation across tasks as a distinct model capability that episodic benchmarks miss and sequential training can improve.
Chat is not available.
Successful Page Load