QCBA: Quantum Coalescence Backdoor Attack against Quantum Convolutional Neural Networks
Abstract
Quantum Convolutional Neural Networks (QCNNs) have emerged as a promising architecture for quantum machine learning in the Noisy Intermediate-Scale Quantum (NISQ) era, using hierarchical convolution and pooling to extract features while progressively reducing the active qubit register. However, this hierarchical structure introduces a distinct challenge for backdoor attacks. Trigger information can be substantially weakened by input downsampling, quantum-state perturbation, and successive qubit reduction, causing attacks developed for conventional fixed-width QNNs to transfer poorly to QCNNs. To address this challenge, we present QCBA, a backdoor attack that directly optimizes the quantum-state distribution at the QCNN readout register. We design a coalescence objective jointly increases the target-class margin of triggered inputs, concentrates their readout states across different source samples, and aligns their centroid with the representation of clean target-class samples. A three-stage training procedure further separates clean pretraining, trigger optimization, and backdoor injection to preserve benign model performance. Experiments on MNIST and F-MNIST across nine convolutional ansatze demonstrate that QCBA achieves attack success rates above 96% and 93%, respectively, while maintaining comparable clean accuracy. QCBA also substantially outperforms representative classical and quantum backdoor attacks. MDS visualization of readout states further shows that QCBA causes triggered representations from different source classes to coalesce with the target-class region, demonstrating the importance of explicitly optimizing the surviving quantum representation for effective backdoor injection in QCNNs.