Shallow ReLU Classifiers on Non-Linearly Separable Data: The Loss Landscape, Edge of Stability, and Beyond
Abstract
Large-step gradient descent (GD) often trains neural networks successfully even when the loss is non-monotone and classical smoothness-based guarantees do not apply. Existing theory, however, largely focuses on linear predictors or linearly separable data, leaving the dynamics of nonlinear models on nonseparable problems less understood. We study a minimal instance consisting of three one-dimensional examples and a two-neuron ReLU network trained with logistic loss. Despite its simplicity, the problem has a nonconvex, nonsmooth objective and exhibits edge-of-stability behavior. We identify a region of parameter space that captures successful classification and reduces the dynamics to a tractable piecewise logistic problem. This perspective allows us to characterize convergence, implicit bias, and a catapult phenomenon under large learning rates. Our results provide a tractable neural network setting for studying gradient descent with large step sizes beyond linear predictors and separable data.