MARL-based System Optimization for Wildfire Sensing Networks
Abstract
Terrestrial-based systems for early wildfire detection often use cameras deployed in remote environments, where images from these cameras are routed through the network and processed by ML models. However, determining an optimal policy for how to route and process collected images is challenging due to highly variable energy availability and network capacity at the edge. Existing methods typically adopt static approaches for determining such policies. They also rely on centralized execution, introducing a potential single point of failure. Multi-agent reinforcement learning (MARL) provides an attractive alternative. However, learned policy performance for MARL-based solutions is dependent on the quality of reward modeling, environment diversity, re-start strategies, and reward credit assignment, all of which can be difficult to define and if done poorly can lead to inconsistent and sub-optimal performance in live deployments. To address these challenges, we propose MARL-based System Optimization for Wildfire Sensing Networks (MARL-OptWSN), a framework for optimizing highly variable image-based wildfire sensing networks. MARL-OptWSN uses a custom training regime built around a Gurobi-based oracle, while leveraging decentralized execution to improve resilience and adapt to changing environmental conditions. Our training approach combines progressive training, oracle guidance, and an agent-specific blame system that identifies decisions responsible for poor outcomes. Empirical experiments show that our approach achieves near-oracle performance with substantially lower decision time, while retaining resilience under highly variable and constrained conditions that negatively affect centralized approaches.