Separating Hazard from Reporting: Rainfall as an Instrument for Under-Observation in Urban Flood Adaptation
Abstract
Cities increasingly prioritize flood adaptation spending using hazard signals built from resident service requests, but a self-selected reporting process produces those requests, so a complaint count confounds how much a neighborhood floods with how likely residents are to call. Existing analyses regress complaint density on demographic co-variates without estimating the latent hazard or stating what is recoverable. We propose to separate the two using precipitation as exogenous variation, and to release the resulting benchmark publicly. Across 316,673 Pittsburgh 311 records, stormwater reporting falls 2.2x from the lowest to the highest poverty quartile, reporting channel varies with racial composition (rho = -0.599), and stormwater requests rise 2.2-2.7x with rainfall while non-weather requests show no response, which is the falsification test any such instrument must pass. The work runs in three phases: a public panel, a hierarchical estimator returning partial-identification bounds on propensity, and a budget-constrained allocation run on raw against corrected signals. The release gives geospatial machine learning a decision-evaluated task on community data.