Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis
为解决单步逆合成预测多样性不足的问题,通过Top-K提示法和大规模数据集训练化学合理性感知的语言模型,结合精细调整与奖励机制提高性能。
为解决单步逆合成预测多样性不足的问题,通过Top-K提示法和大规模数据集训练化学合理性感知的语言模型,结合精细调整与奖励机制提高性能。
This study investigates whether general-purpose large language models (LLMs) can effectively handle complex three-dimensional spatial constraints—such as protein binding pockets, pharmacophores, and anchor fragments—in structure-based drug design to generate plausible binding molecules. To this end, the authors introduce 3D-Fit, the first evaluation framework tailored for multi-spatial-constraint scenarios, which integrates a token-efficient benchmarking strategy with a constraint modeling approach conditioned on heterogeneous spatial information. Experimental results demonstrate that, although current LLMs slightly underperform state-of-the-art diffusion models in generation quality, they already exhibit a promising capacity to jointly incorporate diverse 3D constraints and show strong potential for scalability, thereby validating their prospective utility in structure-based drug design.
Current retrosynthetic planning systems lack a flexible and chemically interpretable evaluation benchmark, making objective performance comparisons difficult. To address this gap, this work proposes the URSA evaluation framework, which introduces— for the first time—a dimension of chemical plausibility that simulates expert chemists’ judgment, complementing formal correctness. URSA enables a comprehensive assessment of retrosynthetic routes by considering reaction feasibility, pathway convergence, and other chemically relevant criteria. The framework evaluates both specialized models and large language models in realistic drug design scenarios, revealing that while the latter show promise in high-level strategic planning, they still significantly lag behind specialized models in task reliability. URSA thus establishes a new, interpretable, and holistic benchmark for the practical evaluation of retrosynthetic methodologies.
This work addresses the limited scientific reasoning capabilities of general-purpose large language models (LLMs) in drug discovery, which cannot be reliably improved merely by scaling up model size or incorporating generic reasoning mechanisms. To overcome this, the authors introduce MMAI Gym for Science—a platform that integrates multimodal molecular data with task-specific training, inference, and evaluation frameworks—and propose a lightweight Liquid Foundation Model architecture designed to effectively model “molecular language.” This approach achieves state-of-the-art or competitive performance across key tasks including molecular optimization, ADMET prediction, retrosynthesis, and drug–target activity prediction, while requiring substantially lower computational resources than both larger general-purpose and specialized models, thereby breaking the performance bottleneck of general LLMs in specialized scientific domains.
This work addresses the challenge of everyday knowledge question answering—including both short-answer and multiple-choice formats—in multilingual and cross-cultural settings. The authors propose a retrieval-augmented generation (RAG) framework that integrates open-source small language models deployed via Ollama, a custom culture-aware knowledge base (CulKBs) combining localized Wikipedia content with real-time DuckDuckGo search results, and iterative prompt engineering. Evaluated on SemEval-2026 Task 7, the system supports English, Spanish, and Chinese while prioritizing user privacy and computational sustainability. All code and resources are publicly released to advance research on lightweight, culturally adaptive multilingual question answering systems.
为解决单步逆合成预测多样性不足的问题,通过Top-K提示法和大规模数据集训练化学合理性感知的语言模型,结合精细调整与奖励机制提高性能。
This study investigates whether general-purpose large language models (LLMs) can effectively handle complex three-dimensional spatial constraints—such as protein binding pockets, pharmacophores, and anchor fragments—in structure-based drug design to generate plausible binding molecules. To this end, the authors introduce 3D-Fit, the first evaluation framework tailored for multi-spatial-constraint scenarios, which integrates a token-efficient benchmarking strategy with a constraint modeling approach conditioned on heterogeneous spatial information. Experimental results demonstrate that, although current LLMs slightly underperform state-of-the-art diffusion models in generation quality, they already exhibit a promising capacity to jointly incorporate diverse 3D constraints and show strong potential for scalability, thereby validating their prospective utility in structure-based drug design.
Current retrosynthetic planning systems lack a flexible and chemically interpretable evaluation benchmark, making objective performance comparisons difficult. To address this gap, this work proposes the URSA evaluation framework, which introduces— for the first time—a dimension of chemical plausibility that simulates expert chemists’ judgment, complementing formal correctness. URSA enables a comprehensive assessment of retrosynthetic routes by considering reaction feasibility, pathway convergence, and other chemically relevant criteria. The framework evaluates both specialized models and large language models in realistic drug design scenarios, revealing that while the latter show promise in high-level strategic planning, they still significantly lag behind specialized models in task reliability. URSA thus establishes a new, interpretable, and holistic benchmark for the practical evaluation of retrosynthetic methodologies.
This work addresses the limited scientific reasoning capabilities of general-purpose large language models (LLMs) in drug discovery, which cannot be reliably improved merely by scaling up model size or incorporating generic reasoning mechanisms. To overcome this, the authors introduce MMAI Gym for Science—a platform that integrates multimodal molecular data with task-specific training, inference, and evaluation frameworks—and propose a lightweight Liquid Foundation Model architecture designed to effectively model “molecular language.” This approach achieves state-of-the-art or competitive performance across key tasks including molecular optimization, ADMET prediction, retrosynthesis, and drug–target activity prediction, while requiring substantially lower computational resources than both larger general-purpose and specialized models, thereby breaking the performance bottleneck of general LLMs in specialized scientific domains.
This work addresses the challenge of everyday knowledge question answering—including both short-answer and multiple-choice formats—in multilingual and cross-cultural settings. The authors propose a retrieval-augmented generation (RAG) framework that integrates open-source small language models deployed via Ollama, a custom culture-aware knowledge base (CulKBs) combining localized Wikipedia content with real-time DuckDuckGo search results, and iterative prompt engineering. Evaluated on SemEval-2026 Task 7, the system supports English, Spanish, and Chinese while prioritizing user privacy and computational sustainability. All code and resources are publicly released to advance research on lightweight, culturally adaptive multilingual question answering systems.