Multi-Agent Reinforcement Learning for Global Climate Policy Cooperation
Abstract
International climate policy requires coordination among countries whose mitigation decisions impose local costs but generate global benefits. Classical game-theoretic and integrated assessment approaches can study these incentives, but typically rely on prescribed behavioural assumptions or equilibrium concepts. Recent work has therefore explored Multi-Agent Reinforcement Learning as a way to model climate-policy cooperation as an emergent outcome of repeated strategic interaction. We extend an existing MARL climate framework with a multi-stage negotiation mechanism in which agents propose agreements, accept or reject commitments, and face penalties for non-compliance. This allows us to study how agreement design affects learned mitigation behaviour, compliance, and climate outcomes. We find that non-binding agreements do little to overcome free-riding, while moderate enforcement improves cooperation and reduces warming. However, excessively strong penalties discourage ambitious commitments, revealing a trade-off between enforcement and participation.