Sequential Local Operator Alignment for Training-Free Model Merging
Abstract
Training-free model merging seeks to combine multiple fine-tuned models into a single model without further optimization on labeled data. Existing methods typically merge individual layers, while overlooking the functional structure of coupled layers underlying the transformer architecture. In the attention mechanism, for instance, the logits are determined by the query-key operator, and the remaining computation depends on the value-output operator. Independently merging query, key, value, and output is sensitive to parameterization ambiguity and can fail to preserve the input-output mapping of the attention head. In this paper, we introduce Sequential Local Operator Alignment, a training-free framework that merges transformer models by sequentially aligning local functional operators rather than individual layers. Our data-aware method uses calibration sets to estimate the local behavior of each functional component, aligns operators sequentially under the intermediate activation of the partially merged model, and then factorizes the merged operators back into valid transformer parameters. Our sequential formulation reduces error accumulation across layers, while the operator factorization step enables rank expansion as a principled mechanism for increasing multi-task capacity. Across CLIP vision benchmarks and RoBERTa models, our approach achieves state-of-the-art training-free merging performance and outperforms competing baselines.