Uncertainty-aware reinforcement learning for chemical language models
This work addresses a critical limitation in conventional reinforcement learning for molecular generation: the neglect of uncertainty in property prediction, which often leads to unstable optimization and the proposal of spuriously high-scoring molecules. To remedy this, the study introduces the first reinforcement learning framework that explicitly incorporates predictive uncertainty into a chemical language model via a dual-path integration mechanism. Specifically, uncertainty is treated both as an auxiliary optimization objective to balance performance and reliability, and as a modulation signal for policy updates to downweight unreliable samples. By integrating ChemProp, random forests, and conformal prediction techniques, the proposed method maintains competitive molecular scores while substantially improving empirical validity—raising the true hit rate from 0.5 to 0.75 and nearly doubling the number of effective molecules generated.