Data Leakage in PPG-based Non-invasive Blood Glucose Prediction
Junyi Shen ⋅ hongmei liu ⋅ Veronica Liesaputra ⋅ Yuan Zhang ⋅ Zhiyi Huang
Abstract
Non-invasive blood glucose level (BGL) estimation from photoplethysmography (PPG) has attracted increasing interest, with numerous machine learning and deep learning methods reporting promising performance. However, inconsistent data-splitting strategies may introduce sample dependence and data leakage, leading to overly optimistic estimates of model generalization. This study systematically evaluates the impact of evaluation protocols on PPG-based BGL estimation using three datasets and three representative models, including TinyML, CBHFF, and CatBoost. We compare random train-test splitting (RTTS), subject-wise RTTS (STTS), participant-aware splitting (PAS), leave-some-participants-out (LSPO), and out-of-distribution (OOD) evaluation. Results show that RTTS generally produces the strongest performance, while performance substantially decreases under subject-separated and cross-dataset evaluation. Notably, PAS did not consistently outperform LSPO, indicating that prior exposure to the same participant did not reliably improve prediction performance. Cross-dataset OOD evaluation showed a further decline in generalization, with most settings producing near-zero or negative $R^2$ values. In contrast, Clarke Error Grid results remained consistently high, with 97-100\% of predictions falling within Zones A+B even when regression performance was poor. This mismatch suggests that favorable clinical error-grid results alone may overstate model reliability. Overall, our findings highlight the need for leakage-aware, subject-independent, and cross-dataset evaluation to obtain a more realistic assessment of PPG-based glucose estimation models.
Chat is not available.
Successful Page Load