Symmetry Alignment Across Depth in Transformer Merging
Krish Malik ⋅ Tasmayu Swain ⋅ Krishna Mudgal
Abstract
Symmetry-based alignment can yield zero-barrier paths between independently trained Transformers at fixed depth. We test whether this picture extends across depth by padding the shallower model with exact identity blocks. The padded model shares the deep model's $L$-block parameter space but lies in an invariant subset where selected residual branches vanish. Across five depth pairs on WikiText-2, Penn Treebank, and Tiny Shakespeare, inherited permutation-and-orthogonal (P+O) alignment removes $90.1$--$94.6\%$ of the naive barrier yet leaves a positive residual of $0.098$--$0.244$ nats in every tested configuration. Additional width ($d\in\{256,512,768\}$ where available) and layer-placement ablations show that the residual depends strongly on representation width and on where identity blocks are inserted. Raw identity-slot count orders the residual within fixed depth $L$ but not across it, and matched-ratio depth trends are corpus-dependent rather than universal.
Chat is not available.
Successful Page Load