xHC: Expanded Hyper-Connections
Xiangdong Zhang ⋅ Xiaohan Qin ⋅ Tuo Dai ⋅ Xiaoming Shi ⋅ Huaijin Wu ⋅ Yebin Yang ⋅ Zhuo Xia ⋅ Shaofeng Zhang ⋅ Yu Wang ⋅ Yu Cheng ⋅ Junchi Yan
Abstract
Hyper-Connections (HC) extend the single residual stream of Transformers into $N$ parallel streams, improving language model pre-training with modest additional FLOPs. Manifold-Constrained HC (mHC) stabilizes this formulation at scale by constraining residual mixing to doubly stochastic matrices. Large gains from $N{=}1$ to $N{=}4$ suggest residual-stream expansion as a promising scaling axis, yet existing HC-family methods typically stop at $N{=}4$. Our experiments on mHC reveal why: directly scaling $N$ beyond 4 yields rapidly diminishing returns, as gains saturate while training compute grows sharply. We trace this to two bottlenecks: fixed-dimensional layer outputs cannot supply enough write-back information for more streams; meanwhile, generating the residual-mixing matrix makes the dominant cost scale cubically with $N$. To address both bottlenecks, we propose **xHC** (E**x**panded **H**yper-**C**onnections), the first HC-family method to achieve meaningful expansion beyond $N{=}4$. xHC enriches write-back signals with local contextual features along the token sequence and uses a sparse residual-stream architecture that updates only $k$ out of $N$ streams while preserving dense access to all streams. In language model pre-training, xHC with $N{=}16$ and $k{=}4$ substantially outperforms both mHC and the vanilla residual baseline. On a 69B MoE model, xHC improves the average downstream score by 5.7 points over mHC (vs.\ +1.3 for mHC over vanilla) with only 2.4\% additional training FLOPs over the vanilla baseline. Scaling law experiments show that xHC matches the loss of a vanilla baseline trained with $1.50\times$ the compute, and its continued improvement with larger $N$ makes expansion rate a practical scaling axis for HC-family models.
Chat is not available.
Successful Page Load