Replay-buffer engineering for noise-robust quantum circuit optimization
Akash Kundu ⋅ Sebastian Feld
Abstract
Deep reinforcement learning (RL) for quantum circuit optimization faces three fundamental bottlenecks: replay buffers that ignore the reliability of temporal-difference (TD) targets, curriculum-based architecture search that triggers a full quantum-classical evaluation at every environment step, and the routine discard of noiseless trajectories when retraining under hardware noise. We address all three by treating the replay buffer as a primary algorithmic lever. We introduce ReaPER$+$, an annealed replay rule transitioning from TD error-driven prioritization to reliability-aware sampling as value estimates mature, achieving up to 4$\times$ improved sample efficiency over the strongest off-policy baselines (fixed PER, ReaPER, and uniform replay) and matching prior on-policy solution quality at up to 32$\times$ fewer interactions. At 12-qubit scale, ReaPER achieves the lowest energy error in the fewest steps which is consistent with ReaPER$+$'s annealing direction, while PER and Vanilla replay find more compact circuits at the cost of higher energy error. Validation on LunarLander-v3 confirms the annealing principle is domain-agnostic obtaining +9% AUC over PER and fixed ReaPER. We further introduce OptCRLQAS, which amortizes expensive quantum-classical evaluations over multiple architectural edits, cutting wall-clock time per episode by up to 67.5% on a 12-qubit task without degrading solution quality. Finally, a lightweight replay-buffer transfer scheme warm-starts noisy-setting learning from noiseless trajectories, without network-weight transfer or $\epsilon$-greedy pretraining, reducing steps to chemical accuracy by up to 85-90% and final energy error by up to 90% over from-scratch baselines, with transfer advantages that grow with system size. Together, these results establish that experience storage, sampling, and transfer are decisive levers for scalable, noise-robust quantum circuit optimization. The anonymized repository is available at : https://anonymous.4open.science/r/annealed-replay-buffer-transfer-for-quantum-optimization-6861/.
Chat is not available.
Successful Page Load