Signal-Driven Fusion for Stable Federated Vision-Language Model Personalization
Abstract
Personalizing vision language models on-device under federated constraints means combining a shared global representation with a locally adapted one, and most approaches do this through a fixed mixing weight chosen once, before deployment. We show this design is unstable in a specific and avoidable way: personalization accuracy peaks early in training and then degrades steadily for the rest of it, no matter how carefully the weight is set. We trace this to client heterogeneity, the initial unreliability of a freshly initialized local adapter, and the fact that a fixed weight simply cannot track how training progress changes over time. DG-BLEND replaces this fixed weight with two lightweight, dynamic mechanisms built for different deployment constraints. The first reads the mixing weight directly from the model's own prediction confidence, adding no parameters and needing nothing beyond what a normal local forward pass already produces, which makes it a useful zero-cost option and a probe of how cheap on-device signals behave under overfitting. The second combines each client's label distribution with a live measure of local-adapter divergence, with the combining rule itself learned through three scalar parameters broadcast from the server, removing the need for any per-client hyperparameter search. Across three datasets (frozen CLIP ViT-B/32, ten clients, strongly non-IID splits, three seeds), the learned variant matches fixed-weight accuracy within 0.65 points while cutting post-peak degradation by 52\%, adding only three scalars of server-side overhead and no extra client-side compute or communication.