CRAE-GUARD: Clean-Referenced Absolute Evidence for Sampling-Robust Quantum Backdoor Detection
Abstract
Quantum neural network (QNN) models can be compromised by backdoor attacks that preserve normal behavior on clean inputs while producing attacker-chosen predictions when a trigger is present. Although recent defenses have begun to detect such hidden behaviors, their reliability remains largely unexplored. We uncover a previously overlooked failure mode: existing model-level detectors can assign different clean/backdoored verdicts to the same trained model when only the finite set of trusted reference examples is changed. Increasing the reference budget alone does not consistently restore both detection accuracy and decision consistency. Our analysis traces the problem to insufficient clean-calibrated decision margin, which allows ordinary sample variation to move a model across the detection threshold. Motivated by this observation, we propose CRAE-Guard, a sampling-robust model-level backdoor detector that separates reverse-trigger recovery from verification and evaluates recovered triggers against clean class-conditional reduced-state references. This clean-referenced evidence produces substantially larger clean-calibrated decision margins and, consequently, more consistent decisions across reference draws. On a fixed QCNN benchmark, CRAE-Guard detects 47--48 of 48 backdoored models across five draws and reduces the model-level flip rate from 0.542 to 0.042. The same qualitative improvement persists on a fixed-width variational quantum circuit, while complementary experiments with QSentry show that sample-draw sensitivity also arises in sample-level detection. These results establish repeated-draw consistency and clean-calibrated decision margin as important reliability dimensions for quantum backdoor detection.