When Error Mitigation Makes Things Worse: Budget-Aware Evaluation, Extrapolation Failure, and the Calibration Trust Boundary
Abstract
Error mitigation is expected to turn noisy quantum measurements into more useful estimates. We show that it can instead amplify error. Under control-offset miscalibration, zero-noise extrapolation (ZNE) is worse than no mitigation on 38-63% of simulated instances and produces extreme tail errors. A small IBM Heron study reproduces the pattern on 80–88% of instances, with mean error 3.0-4.1 times above the raw estimate. The median changes little, so median-only monitoring misses the failures. We then compare methods at the same online shot budget per evaluation and report offline training cost separately. After that cost is amortized, a budget-conditioned neural corrector occupies the low-budget end of the accuracy-cost frontier and falls back to the raw estimate when an ensemble disagrees. Finally, we treat provider-reported calibration as an input that is not bound to the execution-time device state. A projected-gradient attack on this metadata increases the corrector's error 8.1 times within the feature ranges used for training and evades the disagreement monitor on half of the worsened cases. Together, the results connect budget-aware evaluation, tail risk, and calibration integrity in near-term quantum learning pipelines.