The Downstream Cost of Binarizing Forest-Loss Forecasts for Conservation Prioritization
Abstract
Forest-loss forecasts are often trained to predict whether loss occurs, while conservation decisions under limited budgets may depend on how much forest and carbon are threatened. We test whether this mismatch in target formulation changes downstream conservation prioritization. Across temporally separated 1-, 3-, 5-, and 10-year forecasting tasks in Acre, Brazil, matched continuous-target XGBoost and Random Forest models capture 6.5–26.6 percentage points more realized threatened carbon than their binary any-loss counterparts at a fixed 5% forest-area budget. Occurrence-based ROC-AUC can nevertheless favor the worse downstream model, while evaluation against loss-severity labels better recovers the portfolio ordering. In a pre-specified test in Pará, binary models achieve higher any-loss ROC-AUC in all 16 model-family×horizon comparisons, yet continuous targets produce better downstream portfolios in all 12 multi-year comparisons, with smaller effects than in Acre. Together, these results show that target definition is a consequential modeling choice when forecasts feed constrained allocation: prediction targets and evaluation should reflect the quantity the downstream decision actually values.