MLIP Embeddings Predict Extraordinary Materials
Abstract
Extraordinary materials have properties beyond the range of typical training datasets, so discovering them requires predictors that can extrapolate past the labels they were trained on. Here, we test whether the representations learned by machine learning interatomic potentials (MLIPs) support this out-of-support prediction. On seven crystal properties, we hold out the largest 5% of labels and compare a pretrained MLIP, probed or finetuned, against training from scratch and against two classical descriptors under the same predictor. Physics-based descriptors are the conventional choice for extrapolation and led earlier tests, yet we find that MLIP embeddings lead on OOS error on every task. Probed, they correctly identify a median 55% of the held-out extraordinary materials, against 9% from scratch and 23% for the better descriptor, and their error falls faster with data. Finetuning aligns the property along the first principal component of the embeddings, and the pretraining prior extends that axis only a finite distance past the training range: within it the held-out materials stay ordered, beyond it they scatter into the training mass.