EvoLedger: Evidence-Preserving Test-Time Evolution for LLM Agents
Abstract
Frozen Large Language Model (LLM) agents can improve their test-time performance without parameter updates by continu- ally evolving external harnesses, including prompts, memory, and tool-use routines. However, noise, sparse feedback, state de- pendence, and conflicting evidence in long-horizon trajectories can cause unreliable experiences to be repeatedly reused, vali- dated knowledge to be lost during updates, and critical progress signals to be obscured by redundant context. We therefore for- mulate test-time harness evolution as an evidence-management problem and propose EvoLedger, an evidence-preserving evo- lution framework for frozen LLM agents. By coordinating memory use and harness updates according to evidence re- liability and applicability, EvoLedger preserves validated knowledge while distilling long trajectories into information that is genuinely useful for future decisions, enabling more stable cross-episode evolution. Across five benchmark suites and three backbone LLMs, EvoLedger achieves the highest cross-benchmark average performance. On Qwen3-30B-A3B, Gemma-3-27B-it, and GLM-4-32B-0414, it improves the av- erage score over the no-learning baseline by 0.1950, 0.1945, and 0.1930, respectively, and over the strongest automated op- timization baseline by 0.0996, 0.0968, and 0.1097. Ablation studies further demonstrate the complementary benefits of its key designs.