Beyond Reduced-Order Control: Physics-Guided Nine-DOF Rollouts for Wave-Energy RL
Abstract
Wave energy is a substantial renewable resource, but commercial devices must extract energy competitively under changing waves. For the coupled D-Spar con- verter this creates two linked problems: predicting nonlinear device motion ac- curately and choosing power-take-off (PTO) damping quickly enough for con- trol. High-fidelity numerical models retain coupled motion, radiation memory and nonlinear forces, but are expensive to evaluate repeatedly. Rather than further reducing the device dynamics, we learn the expensive nonlinear-force correction inside the full nine-degree-of-freedom model. The remaining excitation, radia- tion, mechanics, PTO and RK4 calculations stay explicit. On development data, a free-recursive 10-s forecast obtains 0.382% macro NRMSE over all displace- ment and velocity channels. A rank-5 proper orthogonal decomposition (POD5) decoder then reconstructs a 15-action damping–energy curve from only three roll- outs, with 0.202% RMS-relative curve error, 100/110 exact sampled optima, exact- or-adjacent selection in every case and 0.0596% maximum regret. The complete measured K3+POD5 controller is not yet 10-Hz real time. Residual RL benefits from POD context but remains 0.606% below greedy. By contrast, exact two-wave sequence MPC gains 4.21% over matched wave-synchronous greedy and 3.57% over canonical 10-Hz greedy across three development seeds, while losing on one seed. These perfect-preview results motivate learning a compact wave-level policy from the MPC value landscape; they do not yet establish RL superiority, unseen- seed performance or real-time deployment.