From Confidence to Explanations: Evaluating Error Detection in EEG Foundation Models Under Distribution Shift
Abstract
EEG foundation models are increasingly used for downstream clinical prediction, yet their reliability under distribution shift remains poorly understood. This issue is particularly relevant to South Asia, where commonly used EEG datasets provide limited representation of local populations and recording environments. A model may retain useful predictive performance across cohorts while the signals used to identify its incorrect predictions fail to transfer. We investigate this distinction using LaBraM and CBraMod for five-class sleep staging. Sleep-EDF Expanded is the source domain, with zero-shot evaluation on ISRUC-Sleep-II, ISRUC-Sleep-III, and UCDDB. We evaluate four error-predicting signals: calibrated confidence, temporal consistency, representation familiarity, and explanation drift. Failure detectors are fitted on source-validation data and evaluated unchanged on the held-out source test set and external cohorts. Confidence provides a strong baseline, achieving source-domain AUROCs of 0.7411 and 0.7551 for LaBraM and CBraMod, respectively, and remains informative externally. Temporal consistency adds significant source-domain gains of +0.0296 and +0.0299 AUROC, but this advantage does not transfer significantly to any external cohort. Representation familiarity behaves heterogeneously, while post-hoc explanation drift adds no statistically supported information beyond simpler signals. These findings show that error-predicting power is itself distribution-dependent. Reliability transfer should therefore be evaluated alongside predictive transfer, particularly for clinical deployment across heterogeneous populations. This is especially relevant to South Asia, where commonly used EEG datasets provide limited representation of local populations and where more rigorous cross-cohort reliability evaluation is needed. Building EEG foundation models with broader South Asian pretraining data, together with systematic reliability audits across institutions, may be important for establishing robust clinical deployment.