Trust-region constraints govern improvement in a closed-loop protein engineering system
Abstract
Closed-loop design systems are now common in the experimental sciences, but few of them report which part of the loop produced the improvement. We report such a system to increase the activity of large serine recombinases (LSRs), which catalyze unidirectional, site-specific integration of multi-kilobase DNA cargoes, making them attractive for therapeutic gene insertion. Each round aligns a 3B-parameter masked protein language model to activity data with a learn-to-rank objective, draws multi-mutant variants by Gibbs MCMC under a soft trust region on mutational distance, assays them by in vitro transcription-translation, and returns the measurements. Across two campaigns, five rounds and 392 assayed designs, the best variant carries five substitutions and reads 165 times parent activity on the same plate. We then ask which component was responsible. Of the components we tested, only the trust region has an effect we can measure: distance to the previously measured set is inversely correlated with activity, absorbing the whole difference between policy-update schedules. We report these findings as a template for future closed-loop campaigns.