OrchestraRL: Learning to Orchestrate LLM Agent Swarms with Entropy-Aware Communication Control
Runze Fan ⋅ Shiyi Jiang ⋅ Junsheng Wang ⋅ Shuangjun Xie ⋅ Xiaodan Li ⋅ Yong Li
Abstract
Multi-agent LLM systems built on frontier backbones often fail to improve reliably as the swarm grows: under strong backbone models, performance frequently plateaus or declines beyond a small number of agents. In a budget-matched controlled intervention on SWE-bench Verified ($N{=}8$, Claude Opus 4.5 backbone, 64-call budget per instance), varying only the communication policy moves resolve rate from 80.4\% under random broadcast coordination (below the 80.9\% single-shot baseline) to 85.5\% under a learned orchestrator, a 5.1-point spread under matched compute. We introduce OrchestraRL, an orchestration layer trained over frozen LLMs. A centralized Macro-Orchestrator reads the swarm's context entropy $H(t)$, computed from the spread of agent output embeddings, and outputs communication mode (via learned entropy thresholds) and topology; per-agent Micro-Coordinators decide whom each agent addresses, when to speak, and at what granularity. Both are trained with REINFORCE and a learned value baseline. Across four benchmarks (SWE-bench Verified, GPQA Diamond, LiveCodeBench v6, SimpleQA), OrchestraRL is comparable to the strongest baseline at $N{=}2$ and achieves the best resolve rate at $N \in \{4, 8\}$ on every benchmark, with the gap widening as $N$ grows. The learned policy recovers interpretable task-dependent communication patterns, including clustered exploration, star-based synthesis, and pipeline reasoning, without these structures being hand-coded.
Chat is not available.
Successful Page Load