Natural Language Actor-Critic: Scalable Off-Policy Learning in Language Space
Abstract
Large language model (LLM) agents---LLMs that dynamically interact with an environment over long horizons---have become an increasingly important area of research, enabling automation in complex tasks involving tool-use, web browsing, and dialogue with people. In these settings, traditional policy gradient methods may suffer from unstable learning and poor sample-complexity due to poor credit assignment. Meanwhile, actor-critic methods address long horizons by learning a critic that provides more granular feedback and enables off-policy (and potentially offline) learning, but are heavily dependent on the fidelity of the critic. In this paper, we propose \emph{Natural Language Actor-Critic} (NLAC), a novel actor-critic algorithm that trains LLM policies using a generative LLM critic that produces values in natural language space. While natural language values have been leveraged in the past to provide a more flexible and actionable training signal, our work is the first to perform general and scalable value learning without relying on in-context aggregation of information from on-policy rollouts. This means our approach can be trained off-policy without policy gradients, offering a more sample-efficient alternative to existing methods. We present results on a mixture of reasoning, web browsing, and tool-use with dialogue tasks, demonstrating that NLAC shows promise in outperforming existing training approaches and offers a more scalable and stable training paradigm for LLM agents.