FedIBS: Federated Vision-Language Adaptation via Intrinsic Bias Selection
Abstract
Federated adaptation of large vision-language models (VLMs) is appealing for privacy-sensitive applications but must contend with client data heterogeneity, strict communication budgets, and on-device inference constraints. Recent methods predominantly introduce additive adaptation interfaces, such as learnable prompts or adapter modules, which attach extra trainable components to the frozen backbone, inflating both parameter counts and inference latency. We challenge this additive paradigm and ask: can robust federated VLM adaptation be achieved through a compact intrinsic interface alone? We answer this question with \textbf{FedIBS} (Federated Intrinsic Bias Selection), a module-free framework that fine-tunes only the biases in feed-forward network projections. FedIBS operates on existing backbone parameters and introduces zero additional inference cost. To make such compact intrinsic adaptation viable under heterogeneous clients, we further propose a Fisher-guided progressive masking mechanism that disentangles shareable and personalized bias coordinates. Each client retains personalized coordinates locally and communicates only the shareable subset for aggregation, enabling sparse communication while preserving client-specific adaptation. Extensive experiments on diverse datasets and heterogeneity settings demonstrate that FedIBS consistently outperforms recent federated prompt- and adapter-based baselines in both personalization and generalization, all while achieving substantially lower communication cost.