Wildfire Danger Prediction as a Positive–Unlabeled Learning Task
Abstract
Wildfire danger prediction is important for preparedness and risk mitigation, and machine learning offers a promising way to integrate meteorological, vegetation, landscape, and human-activity signals at scale. However, most existing methods formulate the problem as supervised Positive-Negative (PN) classification, implicitly treating observations without a recorded fire as true negatives. This assumption is poorly suited as the absence of an observed fire does not imply the absence of fire danger. We instead formulate wildfire danger prediction as a Positive-Unlabeled (PU) learning problem, where recorded events are treated as reliable positives and remaining observations as unlabeled. Using Mesogeos, a large-scale spatiotemporal datacube designed for Mediterranean wildfire modeling, we develop a sampling strategy consistent with the Selected Completely at Random assumption and an evaluation protocol that does not require ground-truth negatives. We compare five PU methods with a supervised PN baseline under a shared LSTM architecture. We find that PU performance depends strongly on modeling assumptions. Methods relying on class separability can exhibit unstable behavior, including seasonal or all-negative collapse, whereas separability- and prior-free variational methods produce more coherent predictions. Taylor Variational PU learning matches the supervised baseline in fire detection, yields spatially precise danger maps, and improves detection of large fire events without explicit fire-size weighting. These results establish PU learning as a viable alternative to PN supervision for large-scale wildfire danger prediction and suggest a broader framework for Earth Observation problems with incomplete negative labels.