Tabular In-Context Learning for Amortized Species Distribution Modeling
Peter Saba ⋅ Rita Menhem ⋅ Chadi Helwe
Abstract
Assessing how marine species will redistribute under climate change requires habitat estimates for thousands of taxa across multiple emission scenarios, time slices, and ocean basins. Supervised species distribution models impose per-task costs such as feature selection, hyperparameter tuning, and spatial validation that scale with every new combination, so even the largest global frameworks cover only a small fraction of the more than 250,000 marine species. We ask whether that cost can be amortized. A tabular foundation model pretrained once on synthetic data (TabICLv2) takes features related to sighting-intensity prediction and their labels as context and predicts sighting intensity in a single forward pass, with no fitting or tuning per species. We test this on three species in three basins: loggerhead turtle (Mediterranean), basking shark (Northeast Atlantic), and blue whale (Northeast Pacific), using sighting records from the OBIS-SEAMAP database and satellite-derived surface-current maps from the EU's Copernicus Marine service, and compare against a tuned Random Forest (RF) under spatial k-means 5 fold block cross-validation, so that each test region is one the model has never seen. With zero training, TabICLv2's mean AUC over the held-out regions is larger than that of a tuned Random Forest whose hyperparameters are selected by nested cross-validation, on the loggerhead turtle (0.99 vs 0.90), the basking shark (0.85 vs 0.83), and the blue whale (0.92 vs 0.90); in every case the 95\% confidence interval on the paired across-region difference spans zero, while its F1 advantage on the loggerhead turtle is significant ($[+0.012, +0.118]$). It also degrades less on the hardest region for two of the three species (worst-fold AUC TabICLv2 0.96/0.31/0.88 vs RF 0.69/0.34/0.85). In-context learning amortizes species distribution modeling: a new species, basin, or scenario costs one forward pass, not a training run, making climate-scale assessment feasible.
Chat is not available.
Successful Page Load