From Confidence Scores to Decision Consequences: A Decision-Theoretic Approach to Communicating Uncertainty for Atrial-Fibrillation Detection
Abstract
While deep learning models have significant potential to improve treatment outcomes, the risk of poor performance on previously unseen data hinders their deployment in practical settings. Uncertainty quantification (UQ) in deep learning may provide an indication of the trustworthiness of model predictions. However, the practical utility of uncertainties for detecting outputs that may negatively impact clinical workflows is often unclear given common UQ evaluation frameworks do not convey the consequences of acting on a model prediction. We investigate decision-theoretic uncertainty as a mechanism for communicating the consequences of using a model's atrial fibrillation classification prediction. Rather than treating confidence scores as an indicator of accuracy, we express Venn-Abers recalibrated confidence scores as expected costs under a toy decision-cost model inspired by screening pathways for wearable devices. We then evaluate when this decision-theoretic representation supports prediction filtering, how its utility changes with asymmetric error costs, and whether its relationship with realised decision-cost remains reliable under synthetic noise-induced domain shift. Cost-aware filtering generally achieved lower realised decision cost at higher retain fractions when the costs of false-positive and false-negative decisions were highly asymmetric; however, its benefit varies with shift severity. This work supports a distinction between communicating predictive doubt from decision consequences.