What Does a Carbon Score Measure? A Case Study of Three Cities in an Urban Emissions Benchmark
Abstract
Cities without a local emissions inventory rely on downscaled global products such as ODIAC, and several recent urban prediction models train on ODIAC carbon as a prediction target. This paper asks what a score on that target measures. ODIAC builds its urban spatial pattern from two components, nighttime lights and a power-plant database with a 2007 base year. On a released three-city benchmark (4959 tiles of 2 km x 2 km), we measure how much of the carbon score each component explains. A monotone transform of the benchmark's nightlight target recovers 54%, 84% and 111% of what a full image+POI pipeline reaches on the carbon target in Shanghai, Beijing and New York, respectively. Most of the remaining error sits on the tiles that hold injected plant values. In Beijing, 13 of 2310 tiles carry 64% of the test-set squared error, and removing them from the test set moves the reported score from 0.440 to 0.771. Adding the plant values from the ODIAC raster as features improves carbon in Beijing and New York, whose labels carry those values, and does not improve it in Shanghai, whose label mostly omits them. We provide a check that flags such tiles from the raster alone, before any model is trained, and we report one falsified pre-registration.