Correlation-aware reward scalarization for reinforcement learning-guided generative molecular design
Abstract
Generative molecular design is often framed as a reinforcement learning (RL) problem in which a policy proposes candidate molecules and is rewarded according to desired properties. Because these properties cannot be measured during generative optimization, they are instead estimated using predictive models, where the reward is typically constructed from multiple in silico molecular property objectives formulated as a multi-parameter optimization (MPO) scalarization. The choice of weights is an important design decision, yet assigning them is difficult: equal weighting implicitly assumes objectives are independent and equally important, and yet drug-discovery objectives are frequently correlated or antagonistic, while manually specified weights rely on expert assumptions that may not reflect statistical dependencies among objectives. We integrate correlation-aware MPO (caMPO), a data-driven scalarization that down-weights redundant objectives, into a molecular scaffold-based RL generator (LibINVENT), recomputing weights online at every optimization step from the molecules generated so far, allowing the reward to adapt to the evolving correlation structure among property objectives. Across 17 scaffolds derived from a public PDE10A dataset, and comparing to a standard equal-weight approach, correlation-aware reward shaping improved property-space quality and alignment with target property profiles, with no significant impact on chemical novelty for both aggregations, and broadened Pareto coverage under arithmetic mean aggregation. These results suggest correlation-aware scalarization may serve as a simple, generator-agnostic reward-design mechanism for RL in experimental molecular science, enabling objective weights to be inferred from data rather than specified manually.