Do Brain MRI Foundation Models Encode Epilepsy-Relevant Concepts? A Confound-Aware Linear Probing Study of Focal Cortical Dysplasia
Abstract
Linear probing is the standard cheap test of whether a clinical concept is linearly decodable from a frozen foundation model's embedding, and the result is almost always read against chance alone. We argue that chance alone is an insufficient reference point for neuroimaging, and show why on a case where the stakes are concrete. We probe frozen BrainIAC embeddings for focal cortical dysplasia (FCD) type II, a surgically treatable cause of drug-resistant epilepsy, using all 170 subjects of the open Bonn FCD-II cohort. The lesion-presence probe reaches AUC 0.694 +/- 0.028, which exceeds chance but not the confound-only baseline. The same embeddings also decode acquisition era (a scanner/protocol period) at AUC 0.986 and age at AUC 0.733, and both are coupled to the clinical label. A classifier using only the one-bit era tag -- with no image at all -- already reaches AUC 0.723, rising to 0.780 when an age bit is added. Restricting the probe to a single acquisition era eliminates between-era variation without modelling it. This reduces performance to 0.566, with a 95% CI of [0.472, 0.631] and a subject-level permutation p = 0.119 (1000 permutations). Hemispheric lateralization never leaves chance. We argue that clinical concept-probing studies in neuroimaging should routinely assess nuisance-variable decodability and compare against explicit confound-only baselines, and that a dataset can be the best available for a disease and still be unable to support the claim that a model has encoded it.