Three ways a small cardiac speech cohort can look solved: an audit of shortcut learning, evaluation protocol and degenerate baselines in a voice biomarker pipeline
Abstract
Speech foundation models make small clinical voice cohorts appear highly predictive, but reported performance can depend more on study design and evaluation protocol than on disease signal. We audit a completed heart-failure voice pipeline comprising 386 sustained-vowel recordings from 122 speakers, including 72 participants with heart failure with reduced ejection fraction and 50 controls. Each recording is represented by handcrafted acoustic descriptors and wav2vec 2.0 embedding components. Holding the cohort, features, and model families fixed, we find that evaluation choices move ROC AUC from 0.754 to 0.934, a substantially larger range than differences between competing model families under a fixed protocol. Three failure modes explain this instability. First, cases and controls contributed different vowels: vowel identity alone, without acoustic or clinical information, reaches AUC 0.754, showing that much of the pooled discrimination can arise from an acquisition-side shortcut. Second, recording-level cross-validation introduces speaker leakage, although its effect is smaller, reducing AUC from 0.914 to 0.895 for logistic regression and from 0.884 to 0.837 for gradient boosting when speaker grouping is enforced. Third, the clinical-record baseline previously used to assess incremental value is effectively non-discriminative because key variables are missing for 61-95% of records; its AUC of 0.532 makes large NRI and IDI values uninterpretable as evidence for added clinical value from voice. After controlling vowel identity and respecting speaker boundaries, a narrower result remains: on sustained /a/ phonation from 122 speakers, logistic regression reaches AUC 0.922 and gradient boosting 0.934, with similar performance on /e/. Age and sex do not explain this separation in the subset with available demographics, although substantial missingness limits that conclusion. Because cases and controls were collected as separate cohorts, residual confounding by recording conditions also cannot be excluded. We conclude that protocol auditing, nuisance-only baselines, participant-level grouping, and validation of clinical comparators are essential before high performance on small voice biomarker cohorts can be interpreted as evidence of clinically meaningful signal.