Reproducible and Interpretable Patient-Specific Seizure Prediction with Feature and Prototype Fusion
Abstract
Reliable seizure prediction from electroencephalography (EEG) could enable timely warning systems for people with epilepsy, but reported performance is strongly affected by patient-specific data splits, alarm-generation procedures, and post hoc operating-point selection. We present a reproducible and interpretable patient-specific seizure-prediction study using long-term scalp EEG from the CHB-MIT dataset. We compare an EEG-only deep-learning baseline with two extensions: feature fusion, which combines learned EEG representations with engineered signal features, and prototype-aware fusion, which introduces interpretable prototype representations into the prediction pipeline. We performed leave-one-seizure-out evaluation across 16 patients and 85 held-out seizures. For each patient, models were trained and evaluated separately, with the held-out seizure excluded from training. Predictions were transformed into alarms using a fixed protocol consisting of a causal 60-second moving average, a 0.5 decision threshold, and a 30-minute refractory period. We report event-level seizure sensitivity, false-positive rate per hour (FPR/h), and pooled area under the receiver operating characteristic curve (AUC). The EEG-only baseline detected 63/85 seizures, achieving 74.12% sensitivity, 0.651 FPR/h, and pooled AUC of 0.753. Feature fusion detected 66/85 seizures, with 77.65% sensitivity and pooled AUC of 0.768, but increased FPR/h to 0.681. Prototype-aware fusion detected 62/85 seizures, achieving 72.94% sensitivity and pooled AUC of 0.770 while reducing FPR/h to 0.515. These results indicate a practical trade-off: feature fusion provided the highest seizure sensitivity, whereas prototype-aware fusion provided the strongest discrimination and lowest false-alarm burden among the evaluated models. To support transparent evaluation, we saved fold-level predictions and model checkpoints, conducted paired fold-wise comparisons, and performed cohort-level analyses of feature-fusion and prototype-fusion contributions. We further distinguish the prespecified primary alarm protocol from exploratory post hoc operating-point analyses, avoiding replacement of primary results with selected threshold configurations. This work provides an auditable framework for patient-specific seizure prediction that integrates reproducibility, interpretable model design, and clinically meaningful alarm-based evaluation.