Learning hydrological error states for process-model correction
Abstract
End-to-end sequence models can jointly learn temporal aggregation, state representations and prediction, but the resulting hydrological state is often difficult to inspect. We revisit the Warren River residual-correction case of Hammad et al. (2026) and replace their prescribed dry/wet residual states with a task-informed representation learned using a vector-quantized variational autoencoder (VQ-VAE). The VQ-VAE yields both a continuous latent state and a set of recurrent discrete prototypes, and a generalized additive model (GAM) is used to predict the GR4J residual correction. The resulting GR4J–VQ-VAE–GAM model performs in the same range as the published state-dependent dual-LSTM correction. Adding the continuous latent representation to explicit predictors improves the GAM readout, with the clearest benefit at high flows. Discrete prototype identity contributes little additional predictive skill, although the prototypes themselves show distinct hydroclimatic profiles, residual distributions and temporal occupancy. We use this case to examine whether hydrological error state can be represented explicitly, inspected, and used for state-dependent correction.