Preference-Graph Policy Improvement for Noisy Biological Sequence Design
Abstract
We introduce Preference-Graph Policy Improvement (PGPI), a method for learning biological sequence edit policies directly from noisy local comparisons. PGPI uses per-variant measurement uncertainties to construct a confidence-gated graph of local improvements, converts its edges into weighted state--action transitions, and trains an iterative edit policy by confidence-weighted behaviour cloning, without fitting an absolute scalar reward. The construction is motivated by two properties: local differencing is insensitive to smoothly varying absolute bias, and a calibrated confidence gate bounds the expected fraction of misdirected edges, subject to a graph-coverage trade-off. Across seven measured fitness landscapes and four noise models, PGPI matches or improves upon direct surrogate ascent most clearly in long, sparse, and noisy regimes, but direct surrogate ascent remains stronger on short, densely sampled landscapes. A controlled corruption study shows that robustness comes from learning over graph-supported transitions rather than from reward-freeness alone; component ablations identify the hard confidence gate and transition representation as the main source of the effect.