Beyond Predictive Performance: Explaining Machine Learning and Tabular Deep Learning Models for Heart Disease Prediction
Abstract
Data augmentation is widely used to improve learning from small and imbalanced tabular datasets, yet it is often treated as a model-agnostic preprocessing step and evaluated primarily using discrimination metrics. Whether augmentation strategies that appear beneficial during model development also preserve discrimination and probabilistic reliability on independent data remains less well understood. This study investigated the question based on using heart-disease prediction as a testbed for reliable tabular machine learning. We benchmark four classical machine-learning models and seven deep tabular architectures using a development cohort of 647 patients and an independent external cohort of 270 patients. Each model is evaluated with the original training data and six augmentation strategies: SMOTE, SMOTETomek, BorderlineSMOTE, ADASYN, CTGAN, and TVAE. Model-augmentation configurations are selected exclusively within the development cohort using internal-validation ROC-AUC and subsequently evaluated on the untouched external cohort. External assessment includes discrimination and classification metrics together with Brier score, calibration intercept, and calibration slope. SHAP and LIME are used to characterize feature-attribution patterns. External evaluation reveals substantial model-augmentation interactions and shows that no single configuration is optimal across all reliability criteria. Among the internally selected configurations, TabTransformer with SMOTETomek achieves the highest external ROC-AUC (0.8540) and lowest Brier score (0.1563), with a calibration intercept of -0.2024 and slope of 1.1489. However, other configurations perform better on individual calibration criteria: NODE-like with TVAE yields the intercept closest to the ideal value of zero (0.0254), while the baseline (no- augmented) Autoencoder classifier achieves the slope closest to one (0.8853). Among classical models, Logistic Regression with SMOTETomek provides the strongest external discrimination (ROC-AUC 0.8349) and lowest Brier score (0.1686). Feature-attribution analyses consistently identify ST-slope, chest-pain type, oldpeak, and exercise-induced angina among the most influential predictors. These findings show that data augmentation should be treated as a model-dependent component of the tabular learning pipeline rather than a universally beneficial preprocessing step. More broadly, selecting models solely by discrimination can obscure meaningful differences in probabilistic reliability and external generalization. Our results motivate evaluation protocols that combine development-only model selection, independent external validation, and multi-dimensional calibration assessment when benchmarking tabular machine-learning systems for high-stakes prediction.