Rethinking Layer-wise Model Merging through Chain of Merges
Abstract
Model merging has emerged as a simple yet effective method to fuse models without re-training. Existing techniques operate at the level of individual layers, thereby overlooking the inter-layer dependencies inherent in deep networks. We show that this simplification induces distributional mismatches in intermediate activations during merging, as changes applied to early layers fail to propagate to downstream ones. We identify these mismatches as a form of covariate shift that compounds across the network. To address this issue, we propose Chain of Merges (CoM), a novel procedure that sequentially merges weights across layers while accounting for inter-layer interactions. In particular, CoM mitigates covariate shift through a series of regression problems, where input activations are recomputed at each step to reflect the updated representations. Additionally, we introduce a dynamic weighting mechanism that prioritizes task-layer pairs most susceptible to performance degradation, improving robustness during merging. Experiments on standard benchmarks demonstrate that CoM achieves state-of-the-art performance across heterogeneous tasks. Code is available in the supplementary material.