COMPASS: Critique-guided Optimization of MetaPolicies for Agentic Systems
Krishna Sayana ⋅ Ketan Todi ⋅ Ambarish Jash ⋅ Sukhdeep Sodhi
Abstract
As large language models are increasingly deployed as autonomous agents, optimizing their behavior in agentic, tool-use environments remains challenging due to the stochasticity of environment trajectories and associated high reward latencies. Existing RL-based prompt optimization methods typically operate at the instance level, often struggling with the high variance and sparse scalar rewards typical of small, human-curated agentic datasets. To address these limitations, we introduce COMPASS (Critique-guided Optimization of MetaPolicies for Agentic Systems), a sample efficient dual-agent framework that shifts optimization toward task-level metapolicies. To overcome sparse feedback, COMPASS introduces a Critique Experience Buffer that couples scalar rewards with dense textual critiques. This mechanism allows an active ``Prompter'' policy to internalize corrective linguistic patterns and amortize experiential self-reflection directly into its weights. Empirical evaluations demonstrate that COMPASS stabilizes training and yields up to 2.4x faster convergence. On complex reasoning (BBEH) and tool-use ($\tau$-bench) benchmarks, our framework improves success rates from 55% to 90% and 74% to 91%, respectively, while naturally discovering specialized algorithmic heuristics.
Chat is not available.
Successful Page Load