From Confidence to Explanations: Evaluating Reliability of EEG Foundation Models Across Distribution Shifts
Abstract
EEG foundation models are increasingly being used for downstream clinical prediction, yet their reliability under distribution shift remains poorly understood. This problem is important for clinical AI systems intended for heterogeneous populations and recording environments, where predictive performance alone may not reveal when a model is likely to fail. In particular, a model may retain useful predictive performance across cohorts while the signals used to identify its incorrect predictions fail to transfer. We investigate this distinction using two EEG foundation models, LaBraM and CBraMod, for five-class sleep staging. Sleep-EDF Expanded is used as the source domain, with zero-shot evaluation on ISRUC-Sleep-II, ISRUC-Sleep-III, and UCDDB. We evaluate four error-predicting signals: calibrated confidence, temporal consistency, representation familiarity, and explanation drift. Failure detectors are trained on source-validation data and evaluated unchanged on the held-out source test set and external cohorts. Confidence provides a strong baseline for detecting incorrect predictions, achieving source-domain AUROCs of 0.7411 and 0.7551 for LaBraM and CBraMod, respectively, and remains informative on the external cohorts. Temporal consistency provides a significant additional gain in the source domain for both models (+0.0296 and +0.0299 AUROC), but this advantage does not transfer significantly to any external cohort. Representation familiarity shows heterogeneous behavior under shift, including significant negative contributions on several ISRUC cohorts and a significant positive contribution for CBraMod on UCDDB. Post-hoc explanation drift provides no statistically supported improvement beyond above simpler signals. These findings show that error-predicting power is itself distribution-dependent: signals that identify failures reliably in-domain do not necessarily retain the same value or direction of association under external shift. Evaluating reliability transfer alongside predictive transfer is therefore essential for assessing the robustness and clinical trustworthiness of EEG foundation models.