Evolving Verifiers for Principled, Training-Free Open-Ended Self-Improvement
Tianyi Xu ⋅ Huazheng Wang
Abstract
Agent self-improvement requires feedback that distinguishes beneficial changes from regressions, yet open-ended tasks often lack an external verifier. Model-generated critiques and rubrics supply feedback, but agents still need a way to choose feedback that improves their responses and adapt it as their behavior changes. We introduce Unified Co-Evolution (UCoE), a training-free framework for selecting and evolving verifiers under a common driving-capacity objective. UCoE synthesizes candidate verifiers from task requirements and compares their estimated ability to improve response quality through selection. Our analysis connects this capacity to pairwise discrimination and motivates worst-group selection on controlled response contrasts. These contrasts have known local preferences and require no deployment-task outcome labels. The selected verifier guides revisions and selects responses. Changes in the agent's external state can shift its response distribution and expose weaknesses in a previously effective verifier. UCoE uses agent experience to propose new candidates and compares them with the current verifier under the same driving-capacity objective. On Arena-Hard-v2, UCoE obtains tie-adjusted pairwise scores of $57.20\%$ with Qwen3-8B and $59.10\%$ with GPT-5.4 against direct generation. On MultiChallenge, UCoE improves Qwen3-8B hidden-rubric accuracy by $7.81$ percentage points.
Chat is not available.
Successful Page Load