On the Design Space of Discrete Diffusion Online Adaptation for Molecular Optimization
Trevor Chen ⋅ Ariel Dai ⋅ Jason Yang ⋅ Riccardo De Santi ⋅ Daniel Khalil ⋅ Wenda Chu ⋅ Nate Gruver ⋅ Pranav Murugan ⋅ Alexander Goldberg ⋅ Maruan Al-Shedivat ⋅ Yisong Yue
Abstract
We study how to design online optimization loops for molecular optimization via adaptation of pre-trained generative models. At test time, we aim to utilize limited oracle feedback to steer generation toward top-molecules achieving high rewards. This creates a design problem coupled across several dimensions, including which candidates receive oracle evaluations, how observed rewards are utilized for model update, and how to overcome pre-trained model biases–to effectively explore over complex design spaces. Despite recent algorithmic progress on each individual component, it remains unclear how they interact in practice, within real-world online adaptation loops. For instance, top-$K$ optimization objectives may make adaptation overly greedy and thereby reduce exploration, while model-debiasing techniques may be unnecessary when exploration is already induced by exploratory acquisition functions, e.g., Thompson sampling. To address this, we conduct a controlled study on discrete-diffusion molecular optimization, finding that well-designed components remain beneficial when combined, indicating that they tackle complementary issues. Together, these components yield an online fine-tuning recipe that outperforms offline fine-tuning and search-augmented baselines across several small-molecule binding affinity and protein fitness optimization tasks, under equal oracle-call budgets and GPU-hour accounting.
Chat is not available.
Successful Page Load