The Coupling Gap: Jointly Monitoring Model Drift and Reliance Drift in Deployed Clinical AI
Abstract
Clinical decision-support (CDS) systems are among the clearest real-world instances of dynamically coupled human–AI systems: deployed models shape clinician trust and reliance, while clinician behavior shapes the data, workload, and labels the model is subsequently evaluated against. Two mature literatures measure this coupling from opposite sides — longitudinal model calibration and dataset shift on one, automation bias and alert fatigue on the other — and a third, more recent methods literature monitors deployed models while treating clinician response as a confound to be adjusted away. We argue for the complementary move: treating reliance drift as a co-primary monitored quantity reported on the same time axis as model performance. We make this concrete rather than hortatory. We specify a dual-drift audit protocol — fixed absolute score bins, a calibration-in-the-large curve, an acceptance curve, and a reliance-discrimination curve, with a weighted-least-squares trend test, a declared minimum effect size, and an explicit rule for confounding medical intervention (CMI) — and we evaluate it in simulation. The simulation yields three results, one of them negative for our own initial proposal. First, the protocol separates silent over-reliance from trust erosion with a 1.4% empirical false-alarm rate and 75%/96% power at a realistic site volume. Second, CMI alone drives the naive over-reliance alarm from 1.0% to 14.2% as intervention effectiveness rises, and restricting the model curve to the non-adhered stratum removes this (≤3.3% throughout). Third, and against our own initial specification, an acceptance-rate curve is blind to reliance-quality erosion: when acceptance stays flat while clinicians stop discriminating good advice from bad, our acceptance-based flags fire at the null rate, and the reliance-discrimination curve we add detects it in only 43% of runs. We report the protocol together with its measured limits, and note that a single published deployment (TREWS) already logs everything the audit needs.