A Streamflow Forecasting Model Learns Proximity, Not Connectivity, from a River Network
Om Phadke ⋅ Justin Baker
Abstract
Predictive models of physical systems often use known physical networks to define their graphs. In practice, any resulting performance gain is usually interpreted as evidence that the model has exploited the physical relation encoded by the graph's edges. However, this interpretation is rarely tested directly, as isolating the network's role requires varying this graph while keeping the model and its training procedure fixed. We test this interpretation in daily streamflow forecasting, where neural networks now match calibrated hydrological models in predicting each basin's next-day discharge from its own meteorological forcings. Across 183 basins in the eastern United States, we hold a standard multi-basin LSTM and its full training configuration fixed and change only the input coordinate that supplies each basin with spatial information about the others. Encoded as a static descriptor of each basin's position in the graph, the network is inert: it changes Nash–Sutcliffe efficiency (NSE) by only $+0.001$ and is not significant at any seed. Encoded instead through neighboring basins' discharge from the previous day, the same network improves NSE by $+0.035$ when that discharge is observed and by $+0.022$ when it must itself be predicted. Controlled substitutions show that proximity, rather than connectivity, carries this gain. At matched distance, a non-upstream basin is as informative as a true upstream neighbor, and the gain declines steadily with distance. The two nearest gauges consequently outperform the drainage network across the 150 connected basins ($+0.081$ versus $+0.043$; paired difference $+0.033$, weakest-seed $p=1.8\times10^{-7}$). This advantage vanishes when neighbor discharge must be predicted, so the result locates predictive information rather than offering a new forecasting method. The gain is reliable only below median flow and is not reproduced by upstream precipitation. Together, these findings show that better predictions after supplying a physical graph do not, by themselves, establish that the model uses the relation encoded by that graph.
Chat is not available.
Successful Page Load