Evaluating the Effectiveness of Location Encoding Methods for Species Distribution Modelling
Abstract
Understanding the spatial distribution of species is a key question in ecology. A variety of solutions, from statical, traditional machine learning, through to deep learning-based methods have been proposed. These methods take sparse observations for a species of interest as input and aim to generate dense spatial outputs indicating the spatial range of the species. Orthogonal to the choice of statistical model used is the question of what input features are most informative. Thanks to recent developments, the number of options has greatly expanded, with choices spanning different climate variables, learning-based encodings, earth embedding models, or raw transformed geographic coordinates. In this work, we introduce a standardised evaluation framework and quantify the effectiveness of different locations encodings in the context of species range estimation, both at the local and global scale.