Tracing the Formation of Super Weights Through Neural Network Artifacts and Dynamics
Abstract
Super weights are individual parameters whose ablation can catastrophically impair a language model. We investigate the architectural and training conditions responsible for their formation. We survey 30 pretrained models spanning 124M-3B parameters and find that no model using non-gated MLPs develops super weights. We then trace their formation in controlled 124M-parameter Transformers across three pretraining corpora and three random seeds that differ only in their MLP activation function. Models with gated or quadratic activation functions develop substantially greater scalar dominance than non-gated models, despite comparable weight outlier ratios and training loss. The distinguishing property is second-order activation growth, not multiplicative gating per se. Extreme hidden activation outliers consistently precede the concentration of large down-projection weights, pointing to early activation outliers as the driver of weight specialisation. Clamping the down-projection input during training eliminates super weight formation entirely, whereas bounding weights or perturbing individual coordinates merely relocates them. These results establish that super weights are a predictable byproduct of quadratic hidden-feature growth during training, with implications for artifact-aware model editing, quantisation, and pruning diagnostics.