Mind the Handoff: Exploiting Classical-Quantum Gradient Interfaces in Quantum Reinforcement Learning
Abstract
Hybrid quantum reinforcement learning (QRL) centers on the interplay between quantum circuit evaluation and classical optimization. This study examines whether this quantum–classical gradient handoff can be exploited to redirect policy learning without modifying the quantum circuit, environment, or data. We introduce QSignFlip, a targeted attack that reverses selected gradient coordinates during training, steering the policy toward an attacker-chosen action. Experiments on CartPole-v1 and LunarLander-v3 show that target-aware coordinate selection shifts policy behavior more effectively than random flipping, increasing the target action’s frequency by up to 54\%. These findings reveal that manipulating the direction of the optimization signal can systematically bias policy outcomes. In response, we investigate a defense that uses quantum information geometry to verify the integrity of the gradient. Our study identifies the quantum–classical gradient handoff as a critical attack surface in hybrid QRL, broadening the security focus to encompass the classical optimization signals that guide quantum policy learning.