Risk Certification for Clinical LLM Deferral under Latent Failure Regimes
Abstract
Deferral policies for clinical language models can be certified by placing a finite-sample bound on the error rate among accepted answers. We study settings in which this pooled risk obscures two distinct modes of evidence use. If the model’s closed-book answer is correct, misleading evidence can corrupt it (corruption); if the closed-book answer is wrong, correct evidence may fail to rescue it (miss). Distinguishing these regimes requires the reference answer, so regime membership is unavailable at deployment and a deployable policy must satisfy both risk constraints without conditioning on the true regime. Across clinical question-answering streams, changing only the gate’s training composition raises pooled AUC from 0.778 to 0.861 while reducing discrimination in the rarer regime to below chance. In a matched evaluation over 33 model–target cells, pooled certification remains valid for its stated aggregate risk, but its unconditional false-certification rate under the regime-wise criterion exceeds the nominal level in every cell. Simultaneous certification exceeds the nominal level in none of them. Regime-wise validity does not by itself ensure useful coverage: the dual procedure certifies no threshold on any split in 14 of the 33 cells. We examine three factors that affect feasibility: per-regime risk targets and confidence allocation, a fixed-sequence threshold ladder that improves certified coverage, and gate training that accounts for regime imbalance. An oracle with access to the regime at inference attains expected certified coverage of 0.84–0.87, compared with 0.22–0.68 for the best deployable gate. The remaining limitation is calibration sample size for multiple-choice outputs and regime-wise discrimination for open-ended outputs.