Deep Learning under Continuous Distribution Shift: The Non-Stationary NTK and Spectral Tracking SDE for Quantitative Finance
Abstract
Standard deep learning relies on assumptions of i.i.d. data and high signal-to-noise ratio (SNR) that do not hold in quantitative finance, where the data-generating process drifts continuously and the irreducible Bayes error is large relative to the predictable signal. This paper develops a theoretical account of deep learning under these conditions. Modeling temporal non-stationarity as bounded Wasserstein drift and treating the training measure as an exponentially weighted aggregate of past distributions, we obtain a temporal generalization bound and identify an optimal decay rate for historical data. We then study the optimization dynamics in the lazy regime via a Non-Stationary Neural Tangent Kernel (NS-NTK), formulated as a stochastic differential equation (SDE). The resulting Spectral Tracking Lag theorem characterizes the neural network as a biased spectral tracker whose lag on each eigencomponent scales as vi/(\eta\lambdai) , so high-frequency alpha signals concentrated in the spectral tail are systematically underfit while smooth beta components are tracked well. We complement this with an analysis of the gradient covariance under low SNR, showing that vanilla SGD aligns its updates with the directions of maximum noise sensitivity rather than with the predictive signal. From these results we derive two algorithms: Principal Variance Reduction (PVR), a low-rank sketching scheme that filters the dominant noise directions, and Eigen-Adjusting Tracking (EAT), a kernel preconditioner that equalizes the tracking speed across the NTK spectrum. We evaluate both on cryptocurrency limit-order-book data and NASDAQ equities, with MLP and Transformer backbones, and report out-of-sample improvements that survive realistic transaction costs and slippage.