NTK Regression under dual power-law model: Deterministic Equivalents via SDE and PDE Methods
Collin Cranston ⋅ Zhichao Wang ⋅ Todd Kemp
Abstract
In the infinite-width limit, training a neural network (NN) via gradient descent reduces to kernel ridge regression (KRR) on a deterministic kernel, the Neural Tangent Kernel (NTK). While the isotropic-input case of the NTK is well understood, the spectra of real world inputs always exhibit power-law decay, and the alignment between this spectrum and the target signal has been shown to govern the generalization behavior of the model. Motivated by this, we study finite-width NTK regression under a dual power-law dataset, where data $\boldsymbol{x} \in \mathbb{R}^{p}$ is drawn with covariance $\boldsymbol{\Sigma}_{jj} = j^{-\alpha}$ and labels are generated from power-law truth vector $\boldsymbol{\beta}_j = j^{-r}$. We introduce a novel technique combining random matrix theory (RMT) and stochastic matrix calculus to derive deterministic equivalents for random matrix functionals that determine the prediction risk, including a new $\textit{signal-weighted}$ resolvent, which yields an explicit formula for the bias-variance decomposition of NTK regression in the high-dimensional proportional regime. As an application we prove a scaling law for finite-width NTK regression, explicitly characterizing the decay rate of the optimally-tuned excess risk jointly in the data and width resources.
Chat is not available.
Successful Page Load