Adaptive Backdoors in Pipeline-Parallel Fine-Tuning
Abstract
Pipeline-parallel fine-tuning is vulnerable to backdoor attacks by workers that control only an intermediate stage of the model. Prior work constructs a backdoor task vector offline and repeatedly injects it during fine-tuning. However, this fixed vector does not account for the changes induced as the model adapts to the owner's fine-tuning data. In this paper, we propose an online attack that recomputes stage-local updates as fine-tuning progresses. Using reconstructed tokens and observed boundary signals, the worker fits a surrogate prefix to track the changing inputs to its controlled stage. This prefix enables auxiliary updates using the current stage parameters without receiving the owner's prompts or accessing other stages' live parameters. Across two owner starting checkpoints, our attack achieves higher success than the tested fixed task-vector baseline after worker removal and 500 safety-fine-tuning updates. An automated judge classifies triggered responses as unsafe on 85.4–88.4% of requests answered safely without the trigger, while at most one of 742 untriggered responses is judged unsafe in either setting. These findings indicate that online adaptation can support persistent, trigger-dependent failures under the evaluated interface.