Learning From Its Own Critique: Dynamic Rich-Feedback Alignment in Multi-Turn Dialog
Siddharth Sabata ⋅ Bogdan Salyp ⋅ William Baker ⋅ Vincent Wilmet ⋅ Rajath Rao ⋅ Mohd Sadiq ⋅ Jofish Kaye
Abstract
Human--AI coupled systems are hard to study directly: interventions on the training loop cannot be ethically grid-searched on real users, and hidden human states cannot be measured. We study this by training a policy against a persona-conditioned user simulator whose emotional state evolves over the conversation. We then test whether using the simulator's textual critiques as an additional training signal improves or destabilizes learning. In \textsc{RLVER-EN}, an English adaptation of an existing multi-turn conversational environment, we compare Group Relative Policy Optimization (GRPO) with failure-routed self-distillation across four Qwen3 scales (8B--32B). The intervention is capability-conditioned: it measurably harms the 8B policy (mean gap $-2.4$ points) yet reliably helps at 32B ($+1.8$; validation accuracy 0.70 in 50 rather than 75 steps), and a cheap offline critique screen predicts this ordering across all four tiers before any training. Thus, rich feedback in multi-turn alignment is useful only when the self-teacher is capable enough to utilize it.
Chat is not available.
Successful Page Load