Federated Reprogramming Knowledge Distillation with Adaptive Aggregation and Multimodal Supervision for Medical Image Classification
Abstract
The rapid development of medical foundation models has shown great promise for healthcare applications. However, fine-tuning these models for downstream tasks remains challenging due to privacy constraints that limit centralized data collection. Federated learning (FL) offers a privacy-preserving solution, yet it must balance model performance with communication and computation costs under non-IID data distributions. Existing parameter-efficient fine-tuning (PEFT) approaches reduce communication overhead but still require clients to host large foundation models, which is impractical on resource-constrained devices. Conventional knowledge distillation (KD) methods fall short in FL due to misalignment between pre-trained foundation models and specific downstream tasks. To overcome these limitations, we propose Federated Reprogramming Knowledge Distillation, a framework in which a frozen medical foundation model resides on the server while only lightweight student models are trained on clients. A server-side reprogramming module aligns the foundation model's feature space with the downstream task, enabling effective knowledge transfer through CKA-based feature alignment and KL-divergence distillation. To further address data heterogeneity across clients, we integrate an adaptive weight aggregation strategy inspired by task arithmetic: client update vectors are used to dynamically assign higher aggregation weights to clients whose updates better align with the global optimization direction, improving convergence stability without requiring proxy data or compromising privacy. Furthermore, we extend it to a multimodal setting by leveraging LLaVA-generated clinical captions to construct image-text pairs for each dataset. In this setting, a full CLIP model serves as the server-side teacher and a compact TinyCLIP model is deployed on clients, enabling multimodal distillation that jointly exploits visual and textual representations. Experiments on three medical imaging datasets under non-IID conditions demonstrate that the proposed method consistently outperforms federated KD and PEFT baselines, offering a strong trade-off between accuracy, communication efficiency, and computational cost in both unimodal and multimodal scenarios.