The Cleaning Layer: How the Convolution Bias on Task-Irrelevant Positions Inflates Curvature and Gradient Noise
Andreas Horlbeck ⋅ Adrian Krenzer
Abstract
The convolution bias places activations on output positions outside the task-relevant region. This couples bias and weight updates through an off-diagonal Hessian entry, and the condition number and bias-gradient variance of the unmasked model grow as $1/(1-\rho)^2$ with the background fraction $\rho$; properties of the forward pass, not the optimizer. We propose the Cleaning Layer, a zero-parameter module removing background-induced activations at their source. For a 1D regression task, we derive the loss surface, gradient decomposition, Hessian eigenvalues, and convergence rate in closed form. Two mechanisms separate: conditioning governs the convergence rate and stability boundary; gradient-noise amplification sets the stationary loss floor. Cleaning attains a lower stationary loss and a wider range of stable learning rates. Adam mitigates conditioning but not the noise; exact Newton preconditioning closes the loss gap entirely, confirming a Hessian-geometric origin.
Chat is not available.
Successful Page Load