Underscoring the Problem: Why Softpick Fails at Initialization
Aryan Sood ⋅ Jaikaran Singh ⋅ Ishaan Bansal
Abstract
Softmax attention gives every token a nonzero weight, which in trained models concentrates into attention sinks and massive activations that widen the dynamic range low-precision inference must cover. Softpick removes this constraint by rectifying scores, eliminating sinks and lowering hidden-state kurtosis, but its advantage fades at scale, a failure the original work attributes to underscoring and dying heads. We reframe this as a normalization problem. Softpick's denominator splits into positive- and negative-shifted sums $D^+$ and $D^-$, used identically in the forward and backward pass, preventing their roles from being isolated. We separate them into a family of operators that independently choose each denominator. The failure originates at initialization: every layer contains rows where $D^+$ is exactly zero, while near-dead rows produce gradient norms above $10^{12}$ regardless of the backward denominator. Only softpick and a stop-gradient variant, which keeps $D^++D^-$ forward but backpropagates through $D^+$ alone, train from scratch. At 230M parameters the stop-gradient operator matches softpick on quantization, has fewer dead heads, and retrieves passkeys more reliably, trailing only on peak attention-weight kurtosis.
Chat is not available.
Successful Page Load