Constrained sample efficient reinforcement learning for heavy ion accumulation in the Low Energy Ion Ring
Borja Rodriguez Mateos ⋅ Verena Kain ⋅ Michael Schenk
Abstract
The Low Energy Ion Ring (LEIR) at CERN turns long lead-ion pulses from Linac3 into the dense bunches required by the Large Hadron Collider's (LHC) high intensity ion program. It accumulates successive multi-turn injections through six-dimensional phase space painting under electron cooling. The accumulated intensity depends sensitively on nine parameters, and the optimum drifts on every relevant timescale: continuously, as the stripper foil that produces the injected charge state ages, and discontinuously, as the machine is recommissioned every yearly campaign or after critical hardware faults. In this paper, we describe an operational Reinforcement Learning (RL) pipeline deployed on LEIR that treats this drift as the central design constraint and outperforms routine operational performance by more than 10\%. Using four years of archived operational data and system identification between 2023 and 2026, we first quantify the drift: generative topographic maps of the nine-dimensional setting space show that both the sampled region and the high-efficiency region move from year to year, and surrogate models fitted on one campaign transfer to the next with near zero rank correlation. We then show how to choose a low-dimensional observation robust to drift: a joint denoising autoencoder with physics-informed auxiliary heads, benchmarked against variational, quantized, and joint embedding alternatives under measured acquisition noise, compresses $\sim$136\,000 bin longitudinal Schottky spectrograms into nine latents that remain informative across campaigns. Policies are trained on per-year heteroscedastic ensemble digital twins and deployed by warm finetuning on the machine: relocking to a new campaign takes a few hundred machine cycles (1 to 2 dedicated hours of training), $\approx$33$\times$ fewer than training from scratch. Finally, exploration is framed as a constrained Markov Decision Process (MDP) with a learned, per-campaign vacuum risk cost: with a fixed Lagrange multiplier the constrained agent explores an order of magnitude more safely than its unconstrained ancestor.
Chat is not available.
Successful Page Load