Similar Predictive Fit but Different Latent Dynamics: Characterizing Learned Dynamical Structure in Personalized Models of Brain Disorders
Rita Huan-Ting Peng ⋅ Quoc Minh Nhat Bui
Abstract
As AI models move toward clinical decision-making and personalized treatment, understanding $what$ a model learns becomes important beyond predictive accuracy alone. This is particularly relevant in clinical populations, where patients are heterogeneous and latent representations learned by large-scale models remain difficult to interpret. We investigate whether personalized latent dynamics can reveal clinically associated differences even when predictive fit is similar. We use our lightweight CNN--Transformer EEG foundation model, pretrained on the large-scale Temple University EEG Corpus (TUEG), to extract transferable segment-level representations. Using the Temple University Epilepsy Corpus (TUEP) as a clinical testbed, these representations are mapped to a shared latent-state space, and sparse multinomial logistic transition distributions (mLTD) are fit independently to each subject to obtain personalized latent transition-dependency graphs $W_n$. Group analyses use $n{=}198$ labeled subjects (99 epilepsy / 99 non-epilepsy). At $k{=}4$, epilepsy subjects exhibit substantially denser learned dependency structure than non-epilepsy subjects (mean nonzero dependencies: 10.01 vs. 5.67; Mann--Whitney $p{=}1.1\times10^{-7}$), with the same group-level pattern observed at $k{=}6$ (19.90 vs. 13.46; $p{=}5.2\times10^{-5}$). Graph-derived features retain moderate subject-level group discrimination under stratified 5-fold subject-wise cross-validation (AUROC 0.68 at $k{=}4$ and 0.65 at $k{=}6$). In contrast, within-subject held-out log-likelihood is nearly identical between groups at $k{=}4$ ($-0.992$ vs. $-0.991$; $p{=}0.95$), and next-state prediction is similarly matched (AUROC 0.855 vs. 0.861; $p{=}0.54$). Therefore, similar predictive fit does not imply similar learned dynamics: the two groups are comparably predictable under their personalized models while differing substantially in the internal dynamical structure learned by those models. This distinction is important even before individual latent states acquire clinical interpretation, because predictive equivalence alone is insufficient to establish that personalized models have learned equivalent internal dynamics. These findings motivate
complementary evaluation of predictive validity and learned dynamical structure before extending such models toward intervention-aware patient world modeling.
Chat is not available.
Successful Page Load