RAG-Enhanced Collaborative LLM Agents for Drug Discovery

๐Ÿ“… 2025-02-22
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF

career value

201K/year
๐Ÿค– AI Summary
Large language models (LLMs) face significant challenges in drug discovery due to the high heterogeneity and semantic ambiguity of biochemical data, coupled with a scarcity of high-quality annotations; domain-specific fine-tuning is prohibitively expensive, hindering real-time scientific data utilization. To address this, we propose CLADDโ€”a novel multi-LLM collaborative RAG agent architecture specifically designed for biochemical data. CLADD enables molecule-level semantic understanding and reasoning without domain fine-tuning, leveraging dynamic retrieval from biomedical knowledge bases, molecular contextual modeling, and zero-shot cross-task generalization. It effectively resolves key challenges including data heterogeneity, semantic ambiguity, and multi-source evidence integration. Empirical evaluation demonstrates that CLADD significantly outperforms general-purpose LLMs, fine-tuned LLMs, and conventional deep learning models on critical tasks such as target prediction, de novo molecular generation, and ADMET assessment. The system exhibits exceptional flexibility, robustness, and reasoning capability.

Technology Category

Application Category

๐Ÿ“ Abstract
Recent advances in large language models (LLMs) have shown great potential to accelerate drug discovery. However, the specialized nature of biochemical data often necessitates costly domain-specific fine-tuning, posing critical challenges. First, it hinders the application of more flexible general-purpose LLMs in cutting-edge drug discovery tasks. More importantly, it impedes the rapid integration of the vast amounts of scientific data continuously generated through experiments and research. To investigate these challenges, we propose CLADD, a retrieval-augmented generation (RAG)-empowered agentic system tailored to drug discovery tasks. Through the collaboration of multiple LLM agents, CLADD dynamically retrieves information from biomedical knowledge bases, contextualizes query molecules, and integrates relevant evidence to generate responses -- all without the need for domain-specific fine-tuning. Crucially, we tackle key obstacles in applying RAG workflows to biochemical data, including data heterogeneity, ambiguity, and multi-source integration. We demonstrate the flexibility and effectiveness of this framework across a variety of drug discovery tasks, showing that it outperforms general-purpose and domain-specific LLMs as well as traditional deep learning approaches.
Problem

Research questions and friction points this paper is trying to address.

Overcoming domain-specific fine-tuning in drug discovery.
Integrating vast scientific data for drug development.
Addressing data heterogeneity and ambiguity in biochemical tasks.
Innovation

Methods, ideas, or system contributions that make the work stand out.

RAG-enhanced LLM agents
Dynamic biomedical data retrieval
Domain-free drug discovery framework
๐Ÿ”Ž Similar Papers
๐Ÿ’ผ Related Jobs
AI Data Engineer--LLMs / Agentic Systems
Pfizer
The annual base salary for this position ranges from $106,000.00 to $176,600.00. In addition, this position is eligible for participation in Pfizerโ€™s Global Performance Plan with a bonus target of 15.0% of the base salary and eligibility to participate in our share based long term incentive program. We offer comprehensive and generous benefits and programs to help our colleagues lead healthy lives and to support each of lifeโ€™s moments. Benefits offered include a 401(k) plan with Pfizer Matching Contributions and an additional Pfizer Retirement Savings Contribution, paid vacation, holiday and personal days, paid caregiver/parental and medical leave, and health benefits to include medical, prescription drug, dental and vision coverage. Learn more at Pfizer Candidate Site โ€“ U.S. Benefits | (uscandidates.mypfizerbenefits.com). Pfizer compensation structures and benefit packages are aligned based on the location of hire. The United States salary range provided does not apply to Tampa, FL or any location outside of the United States. Relocation assistance may be available based on business needs and/or eligibility.
United States - Massachusetts - Cambridge