Statistical Mixing Guarantees for Contractive Echo State Networks
Abstract
Reservoir computing turns sequence learning into linear regression on fixed dynamical features, but those features come from a single highly dependent trajectory, so the true amount of usable data is often unclear. We develop a transport-based learning theory for contractive echo state networks driven by stochastic inputs by viewing the reservoir as a Markov process and proving a Wasserstein contraction that guarantees a unique stationary state law and explicit mixing rates. Using contraction-to-concentration tools, we obtain finite-sample deviation bounds for time-averaged features and empirical covariance estimates, and translate them into excess-risk and stability guarantees for ridge-trained linear readouts. The theory yields an explicit effective-sample-size principle and exposes a sharp trade-off: making reservoirs more “critical” can increase dynamical memory while slowing mixing and raising the data requirements for reliable training. Overall, the results complement echo-state and capacity analyses by providing verifiable design rules that link stability parameters to mixing time, generalization, and forecasting robustness under independent or weakly dependent inputs.