What Adapter Reuse Buys: Early Gains and the Cost of Source Selection
Abstract
We isolate one decision in continual task acquisition: whether to train a new adapter from the identity initialization or initialize it from a copy of a previously trained adapter. We study this decision for Orthogonal Fine-Tuning (OFT) using a frozen backbone, a library of 23 source adapters, and 12 Super-Natural Instructions tasks spanning question answering, classification, generation, and editing. All tasks are cast as text-to-text prediction, and we score each generated output against its reference using lowercased whitespace-token F1, the harmonic mean of token precision and recall. We compare identity initialization, which we call cold training, with six reuse policies over three paired seeds and eight validation checkpoints. Every reuse policy improves mean validation-curve token F1 by 3.97--5.96 points, with target-bootstrap intervals above zero. The benefit is sharp but short-lived: 58--88\% of the summed gain appears at step 25, and no policy improves mean final-test quality. Random reuse captures 68\% of the best policy's mean curve gain, while a learned five-feature selector has no resolved advantage over simpler rules and fully charged profiling leaves no common wall-time coverage. Thus, OFT reuse provides an early optimization head start in this acquisition setting, but the value of paying to select a particular source remains unestablished.