Safe-Fair MACPO: Burden-Fair Constrained Policy Optimization for Safe Multi-Agent Reinforcement Learning
Ankita Kushwaha ⋅ KIRAN RAVISH ⋅ Preeti ⋅ Pawan Kumar
Abstract
Safe multi-agent reinforcement learning usually constrains team-level safety cost, but a safe team can still assign the same agent to repeatedly wait, yield, detour, absorb intervention, or lose access to a scarce resource. We study this failure mode as *burden unfairness under hard safety*. Safe-Fair multi-agent constrained policy optimization (Safe-Fair MACPO) augments MACPO with agent-wise burden logging, a temporal fairness-debt state, a fairness critic, and a two-cost trust-region update. Safety is lexicographically primary: recovery may override fairness to avoid unsafe actions, while the resulting imbalance is stored as debt and penalized later. We give a safety-first Pareto theory showing that the exact safe-fair constrained selector is strongly Pareto efficient for reward, safety, and fairness, with hard-safety and temporal burden-gap certificates. Empirically, the primary reported rows keep zero hard violations while improving useful burden balance: HalfCheetah-2x3-Fair improves selected safe return from $2145.5$ to $2637.4$ and Jain burden from $0.903$ to $0.977$; Safety-Gym MultiGoal-Fair reaches the lowest matched-budget scalar fairness cost with zero safety cost; and VMAS-SafeMultiGoal improves over recovery-only attribution controls. We separate social/resource fairness from mechanical workload imbalance and reserve mixed or recovery-explained rows for qualified appendix analysis.
Chat is not available.
Successful Page Load