Gradient Clipping on Separable Data: Asymptotic Alignment and Persistent Bias
Mansouri ⋅ Nima Mollaei ⋅ Yongtao Wu
Abstract
Can gradient clipping change the asymptotic classifier selected by training on separable data? It depends. For per-example clipping at the empirical median of the current per-example gradient norms, we construct an exact separable family in which the integer $r\ge2$ controls sample multiplicity. The threshold tends to zero while a rare, maximum-margin-relevant example remains clipped at every positive iterate; duplicating geometrically redundant constraints leaves the Euclidean hard-margin feasible set unchanged but changes the limiting separator, whose normalized-margin ratio is $\sqrt{2/(r^2+1)}\to0$. In contrast, asymptotic updates are aligned with a fixed steepest-descent geometry preserving its max-margin implicit bias under suitable step conditions. Fixed-threshold generalized gradient norm clipping (GGNC) is a concrete positive case: it only rescales the norm-specific steepest-descent direction and then after that it becomes inactive. Experiments reproduce the exact law, show finite-horizon separation under modest perturbations, and match the predicted transient-to-inactive behavior of fixed clipping.
Chat is not available.
Successful Page Load