Policy-Level Exploration for Coordinated Multi-Agent Reinforcement Learning
Abstract
In multi-agent reinforcement learning (MARL), agents must not only discover high-reward behaviours but also reliably reproduce coordinated behaviours over time. A fundamental challenge arises from a mismatch between action-level stochastic exploration and temporally extended coordination. To address this issue, we propose a policy-level exploration framework with Deterministic Action Execution (DAE) for MARL, which shifts locally random action exploration to temporally extended policy selection via an option-based exploration mechanism. Specifically, at each time step, each agent selects a policy from a shared policy set\textemdash comprising one exploitation policy and multiple exploration policies\textemdash using a learned option policy, and actions are executed deterministically conditioned on the selected policy. By maintaining deterministic execution under temporally extended policy selection, DAE facilitates temporally coordinated behaviours across multiple time steps. Comprehensive experiments on four multi-agent tasks\textemdash Predator Prey, StarCraft II micro management challenge (SMAC), SMACv2, Google Research Football, and Multi-Agent Coordination benchmark (MACO)\textemdash demonstrate that DAE outperforms mainstream baselines, achieving superior sample efficiency and final learning performance.