Interpretable Machine Learning from Short Duration Photoplethysmography for Peripheral Artery Disease Detection
Abstract
Machine learning from physiological signals is often dominated by deep models that require long recordings and offer limited interpretability. We develop an interpretable, feature-based machine learning framework for detecting peripheral artery disease (PAD) from only four seconds of photoplethysmography (PPG). In contrast to prior deep learning approaches using longer PPG recordings, our approach explicitly identifies waveform characteristics associated with disease and evaluates whether these physiologically grounded features provide predictive information beyond conventional clinical risk factors. We analyzed 5,237 PPG waveforms from 2,362 unique patients with concurrent ankle brachial index (ABI) measurements (the current clinical gold standard for a PAD diagnosis), representing the largest paired PPG-ABI dataset studied to date. From each waveform, we generated 1,525 candidate features describing morphology, timing, derivatives, symmetry, and other time series characteristics. We performed mutual information-based feature selection independently across two clinical datasets and assessed feature selection stability on each of the two sites using 100 bootstrap iterations, retaining 78 features with consistent informativeness across datasets. We trained a linear support vector machine with balanced class weights and evaluated performance using ten-fold stratified group cross-validation, with patients grouped across folds to ensure evaluation on unseen individuals, and differences in AUC between feature-set configurations assessed with DeLong’s test for correlated ROC curves. Using PPG features alone, the model achieved a ROC AUC of 0.831 (95% CI 0.796-0.866) and precision-recall AUC of 0.697 for ABI-defined PAD. PPG-based models substantially outperformed models using demographics and clinical comorbidities alone. Among the retained PPG features, all 78 had greater mutual information with the prediction target than the 30 evaluated clinical features. Adding smoking status increased ROC AUC to 0.845, while adding the full set of demographic and comorbidity variables to the PPG + smoking status model had no statistically significant improvement. The selected features also provided physiological interpretability. Two of the strongest features, normalized waveform width at half amplitude and maximum rising slope, were associated with ABI and matched known signatures of increased arterial stiffness and dampened distal perfusion. Model performance was also consistent across sex, race, ethnicity, coronary artery disease, diabetes, and end-stage renal disease subgroups, with no subgroup showing statistically different discrimination from the overall cohort. Similar performance across two clinical datasets collected using different equipment further supported robustness to measurement conditions. These results show that structured feature extraction from short physiological signals can support accurate and interpretable machine learning without relying exclusively on opaque representations. More broadly, the study illustrates a template for digital biomarker development that jointly addresses feature selection, prediction, physiological interpretation, subgroup auditing, and cross-site robustness: dimensions that are often treated separately in ML for health.