Learning Flexible Context Management Unlocks Test-Time Scaling
Abstract
Test-time scaling improves model performance by allocating additional compute during inference. For modern agentic tasks whose reasoning horizons exceed LLM context limits, effective scaling requires context management: deciding how fresh context is allocated and how information is reused across context windows. Existing approaches typically prescribe fixed strategies through the agentic harness, even though the best strategy may change as problem solving progresses and the environment evolves. We introduce HERMES, a family of simple, parameterized harnesses that progressively shift context-management decisions from the harness to the model. Using two generic mechanisms, i.e. delegation and digestion, HERMES enables flexible orchestration of fresh contexts over extended reasoning horizons. While frontier models can effectively exploit this flexibility, smaller open-source models exhibit substantial capability gaps. We propose a two-stage learning framework that teaches models to adaptively manage context. Experiments show that our methods improve long horizon test-time scaling and enable smaller models to benefit more effectively from additional inference-time compute.