Fidelity Regret: Verifiable Feedback with a Misspecified Simulator
Noah Flynn
Abstract
Reinforcement learning from verifiable feedback assumes that the verifier evaluates the quantity of interest. Genetic-circuit design violates this assumption: a topological verifier can establish that a program is a valid circuit computing the target function, whereas ranking its physical performance requires a simulator. We define \emph{fidelity regret}, the measured reference quality forfeited when a cheap simulator selects the design, and compute it exactly by exhaustive enumeration over 21{,}362 designs, 5 measured gate libraries, and 3 organisms. Mean normalized regret is 0.369 over 72 cells; 25 cells (34.7\%) have regret at least 0.5, and only 6 have zero regret. The discrepancy depends strongly on the reference axis: the yeast library \texttt{SC1C1G1T1} has mean regret 0.113 for cytometry quality and 0.801 for growth burden, and every evaluated library--circuit-size pair has positive burden regret. Topological verification ties all valid gate assignments, while an oracle upper bound on the per-gate-admissibility verifier class retains positive regret in 30 of 52 cells. Design-space optimization and feedback-guided search with frozen language models reach the high-regret regions; on a matched zero-regret control, additional optimization instead reduces regret. Mitigation is budget- and library-dependent. A fixed rule that spends $\alpha{=}0.34$ of the budget re-scoring top cheap-simulator designs achieves mean excess regret 0.016, compared with 0.305 for cheap-only selection, and improves on cheap-only selection in 330 of 360 cells. More broadly, any workflow optimized through a simulator or digital twin can satisfy its verifier yet perform poorly on the intended physical objective. Fidelity regret provides a transferable template for detecting this gap before optimization and comparing budget-aware remedies across scientific domains.
Chat is not available.
Successful Page Load