Trustworthy Machine Learning in the Real-Time Decoding Loop
Abstract
A surface-code decoder has to process each round of syndrome data about as fast as the quantum processor produces it, so a learned model with variable inference time is risky to place in the decoding loop. We study a design in which a learned sidecar estimates context-conditioned matching weights and publishes them only through a trusted, double-buffered admission shim that clips every weight to a band around its static value, so that model inference never runs on the real-time thread. Using stim circuit-level noise with sparse blossom and Fusion Blossom, we find that decode time is set mainly by syndrome density (R² = 0.995 to 0.997 against device drift), while the loaded weights raise mean decode time by at most 9.6 percent on sparse blossom and by up to 148 percent on the Fusion Blossom software. Thus, we treat backlog and prior staleness as separate budgets. For service that varies with the hardware recalibration cycle, we prove that the backlog stays tight whenever the cycle-mean service under static weights, inflated by the measured weight margin, is below capacity, for every model that publishes through the shim. A Lindley-queue simulation driven by measured decode times changes from stable to divergent within 7 percent of the predicted boundary at four code distances. Under a detector-error-model surrogate of a moving, leakage-like hot column, a logistic-regression localizer trained on five-frame syndrome burn-ins lowers the logical error rate by 12 to 34 percent relative to static weights and recovers 49 to 94 percent of an oracle's gain. However, a stale prior that points at the wrong column is 16 to 18 percent worse than static weights, so prior updates must arrive within a break-even staleness of 0.85 to 1.16 mean dwell times. Under smooth drift, a fresh prior changes the logical error rate by at most 3.1 percent, which suggests that adaptation may not be worth its cost in that regime. All of our experiments are simulations.