Tool-Integrated Reasoning via Hierarchical Multi-Agent Reinforcement Learning
Abstract
Tool-Integrated Reasoning (TIR) enables large language models to solve complex tasks by invoking external tools during reasoning. Existing methods typically rely on either single-agent reinforcement learning for unified optimization or multi-agent systems for explicit role decomposition. However, single-agent RL often entangles high-level planning with low-level execution, while multi-agent systems may suffer from role mismatch when different roles are optimized independently. To address this limitation, we introduce the principle of decoupled inference with hierarchical alignment: high-level planning and low-level execution should remain separated during inference, while their policies should be aligned through shared task feedback during training. We instantiate this principle as HATR, a Hierarchically Aligned framework for Tool Reasoning. At inference time, HATR represents each reasoning turn as a tree-structured decision process, where a high-level planning role determines the action branch and step goal, and a low-level execution role realizes this decision through tool use or internal reasoning. During training, HATR performs hierarchical multi-agent reinforcement learning by sampling online trajectories, deriving preference signals from task-level rewards, and alternately optimizing policies across decision levels. This cross-level optimization aligns roles across the multi-agent hierarchy while preserving their functional boundaries. Experiments on mathematical reasoning and question-answering benchmarks across multiple backbones show that HATR consistently improves task performance over single-agent and multi-agent baselines.