Specialize, Don't Pretrain: Lightweight Adaptation of Vision Foundation Models for Climate Monitoring
Abstract
Remote-sensing imagery differs from web imagery in that the same scene can be observed through RGB, multispectral, and SAR sensors. This sensor diversity has motivated domain-specific remote sensing foundation models, but their pretraining corpora are often narrower than those used by general-purpose vision models. We therefore ask whether remote-sensing structure can be introduced by extending a strong web-pretrained RGB encoder to sensing modalities it has never observed. We address this challenge with two simple objectives: cross-sensor alignment and spectral-index prediction. Starting from DINOv3, this adaptation updates only 2.03\% of its parameters using 40K unlabeled images and 5k co-registered multisensor tiles. Across diverse downstream tasks, including flood and burn-scar segmentation, the resulting encoder achieves the strongest overall performance among the evaluated models. Our results show that lightweight specialization offers a compute-efficient foundation for climate monitoring across remote-sensing modalities and tasks.