PMO-Dock: Benchmarking Docking, Specificity, and Generalization in Molecular Optimization
Abstract
The Practical Molecular Optimization (PMO) benchmark standardized evaluation in molecular optimization, but it is built on simple property-based oracles that do not capture structure-based drug design. As the field has moved to docking-based objectives, evaluation practice has become unstandardized, and different methods report results on different docking tasks, often without a shared protocol for fair comparison or sample-efficiency constraints. To address this, we introduce Practical Molecular Optimization for Docking (PMO-Dock), a benchmark and protocol for docking-based optimization consisting of 25 tasks covering hit generation, lead optimization, and a new specificity task requiring strong on-target binding while penalizing off-target interactions. The protocol separates model development from final evaluation by assigning validation tasks for hyperparameter tuning and strictly held-out test tasks for reporting, with no test-task use during development. It also enforces strict oracle budgets to reflect realistic drug-discovery settings. We benchmark four diverse high-performing methods spanning different optimization paradigms, Saturn (reinforcement learning), GenMol (discrete diffusion), Genetic-guided GFlowNet, and Chemlactica (LLM-based). Our analysis shows no universally best method across tasks, substantial differences in hyperparameter sensitivity, and method-dependent transferability, where some methods benefit from global hyperparameter selection while others require task-local tuning on sufficiently similar validation tasks. This benchmark provides a concrete evaluation standard for measuring progress on generalizable, sample-efficient molecular optimization.