AnchorOpt: Structured Runtime-Control Optimization for Tool-Using Agents
Abstract
Tool-using agents make a sequence of consequential runtime decisions about when to call tools, modify state, retry, or stop. Optimizing this behavior can quickly become combinatorial because interventions may differ in where they act, what information they use, and what corrective actions remain feasible. We introduce AnchorOpt, a decision-centric framework for structured runtime optimization. Given residual failures, AnchorOpt attributes them to runtime decision boundaries. Each boundary jointly determines what a local controller can observe and what interventions remain feasible, pruning invalid choices before evaluation. AnchorOpt further decomposes optimization by sequentially adding controllers against a moving incumbent and alternating between policy optimization under fixed signals and signal expansion when the current representation is exhausted. This converts open-ended agent adaptation into a sequence of smaller, constrained runtime decision problems. Initial experiments on tool-using and memory-intensive agent benchmarks show promising gains and support the value of the proposed structure.