Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements
Zhi Zheng ⋅ Rongsheng Chen ⋅ Yunpeng Ba ⋅ Zhenkun Wang ⋅ Yee Whye Teh ⋅ Wee Sun Lee
Abstract
Reinforcement Learning (RL) has been promising in single-turn LLM fine-tuning. However, long-horizon agentic reasoning introduces increasingly branching interactions and sparse rewards, exposing several RL limitations: heavyweight backpropagation makes it impractical to fine-tune larger LLMs, and longer-horizon trajectories make credit assignment substantially harder. This paper argues that evolution strategies (ES) can be a better choice for fine-tuning long-horizon LLM agents. Compared with agentic RL, ES offers three key advantages: **1)** Model Scalability: ES enables full-parameter optimization requiring only minimal, inference-level GPU memory, making it possible to fine-tune large LLMs. **2)** Flexibility: Its lightweight interface makes ES fine-tuning easy to compose with prompt-space evolution; and **3)** Long-Horizon Scalability: ES avoids decomposing rewards across horizons, yielding better long-horizon scalability than Agentic RL. This paper proposes \ourmethod, a full-parameter agentic fine-tuning framework tailored to flexible parameter--context co-evolution. To improve the exploration--adaptation trade-off, Agentic ESOpt further introduces a cosine decay schedule of the perturbation scale $\sigma$. We evaluate Agentic ESOpt across both train-time fine-tuning and agentic test-time compute settings. On long-horizon Sudoku, Agentic ESOpt outperforms RL methods by 12.50\% with Qwen3.5-4B. Across Math and DocVQA, it improves Qwen3.5-4B over Agentic GRPO, and is applicable to Qwen3-30B-A3B. In test-time automatic heuristic design, Agentic ESOpt performs online prompt--parameter co-evolution, improving its baseline in 28 of 36 settings. Code is available at https://anonymous.4open.science/status/Agentic-ESOpt-anonymous-197C.
Chat is not available.
Successful Page Load