Can an MLP absorb its skip connection exactly?
Antonij Mijoski ⋅ Marko Karbevski
Abstract
The benefits usually attributed to skip connections are optimization-theoretic: a smoother loss landscape and better gradient propagation. We ask a representational question instead: given a residual block x ↦ x + MLP(x), does a residual-free MLP of the same width compute the same function? The answer is no, unconditionally and at every depth, for every activation used in current frontier models (ReLU², ReGLU, SwiGLU, GeGLU). For ungated ReLU and GELU absorption is possible, but only on a set of weights of measure zero. The two families are therefore generically disjoint: removing a skip connection and retraining cannot recover the same function at equal width.
Chat is not available.
Successful Page Load