Verbal GRPO: Reference-Free Contrastive Memory Distillation for Self-Evolving Research Agents
Shujing Dong ⋅ Yuan Ling ⋅ Zhenyu Zhang ⋅ Meng-Chen Wu ⋅ Nataraj Mocherla ⋅ Xuxiang Wu ⋅ Balaji narayanan ⋅ Ayush Goyal
Abstract
Reinforcement learning methods like GRPO effectively align LLM agents through group-relative preference optimization, but require gradient access and costly parameter updates, making continual improvement after deployment infeasible. We propose Verbal GRPO, a framework for deep research agents that performs the functional equivalent of GRPO entirely in context space, at test time: we sample a group of trajectories, rank them via an aspect-based judge panel, and distill the semantic advantage from the best-vs-worst contrast into persistent normative criteria consolidated in $\mathcal{M}_{norm}$. Unlike gradient-based GRPO, Verbal GRPO requires no parameter updates and is API-compatible with black-box models; its monotonic, deduplicated memory accumulates knowledge across tasks without catastrophic forgetting. Unlike concurrent training-free RL methods that rely on ground-truth rewards and are limited to verifiable domains (math, code, web search), Verbal GRPO is reference-free, enabling deep research agents to self-evolve quality standards solely through contrastive self-play ($y^+ \succ y^-$) in open-ended domains where no automated verifier exists. Evaluations on ResearcherBench and ResearchRubrics show Verbal GRPO achieves up to 96.9\% improvement rate, with significant within-method gains across four model backbones ($p < 0.001$, Wilcoxon signed-rank). Frozen criteria generalize to held-out tasks (+0.30) and an independent benchmark (+0.51), while a matched random-criteria control (+0.13 vs.\ +0.84) confirms that improvement depends on the semantic content of the distilled criteria rather than prompt length alone. Verbal GRPO offers a practical, interpretable path toward agents that keep learning after deployment.
Chat is not available.
Successful Page Load