How Reliable Are Deletion and Retention Audits for Cumulative EHR Risk Models?
Iris Szu-Szu Ho ⋅ Bruce Guthrie ⋅ Konrad Rawlik ⋅ Sohan Seth
Abstract
Clinical deployment of EHR risk models requires knowing not only which diagnoses are associated with an outcome, but which recorded evidence the model actually uses. This is challenging because diagnosis histories are cumulative, correlated, and often substitutable. We develop a past-only KEEP/HIDE audit that distinguishes diagnostic sufficiency, whether retained evidence sustains a prediction, from in-context reliance, whether deleting a diagnosis from an otherwise intact history lowers it. Beyond clinical plausibility, we evaluate faithfulness, edited-input detectability, sensitivity to measured case mix, and robustness to alternative temporal edits. We apply the audit to 3.6 million UK records for depression, coronary heart disease (CHD), and type 2 diabetes (T2DM). HIDE achieves higher top-10 comprehensiveness than Gradient $\times$ Input and local occlusion for all outcomes. Depression shows robust reliance on anxiety, alcohol misuse, and substance misuse. CHD relies mainly on cardiovascular and metabolic morbidity, although odds ratios recover eight of HIDE's top ten diagnoses and achieve higher comprehensiveness. T2DM relies broadly on diabetes-related and cardiometabolic evidence, but its leading diagnoses vary under measured case-mix standardisation and alternative temporal edits. Global HIDE preserves cumulative structure and remains closer to observed records than anchored KEEP. These results show when model auditing adds value beyond simple association and when attribution conclusions require sensitivity analyses.
Chat is not available.
Successful Page Load