A Retrieval-Augmented Knowledge Mining Method with Deep Thinking LLMs for Biomedical Research and Clinical Support

📅 2025-03-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
To address challenges in biomedical knowledge graph (KG) construction—including terminological complexity, data heterogeneity, dynamic knowledge evolution, and cross-document reasoning—this paper proposes IP-RAR, an integrated retrieval–reasoning framework. IP-RAR introduces a novel two-stage paradigm combining retrieval with self-reflective, deep reasoning to overcome bottlenecks in multi-hop inference and implicit knowledge recall. Leveraging this framework, we construct BioStrataKG—a dynamically evolving biomedical KG—and BioCDQA, a cross-document question-answering benchmark. We further implement LLM-powered automated KG construction, RAG-enhanced generation, and multi-hop semantic matching. Experiments demonstrate a 20% improvement in document retrieval F1-score and a 25% gain in answer accuracy. The framework effectively supports clinical decision-making for personalized medication and facilitates research trend analysis and gap identification.

Technology Category

Application Category

📝 Abstract
Knowledge graphs and large language models (LLMs) are key tools for biomedical knowledge integration and reasoning, facilitating structured organization of scientific articles and discovery of complex semantic relationships. However, current methods face challenges: knowledge graph construction is limited by complex terminology, data heterogeneity, and rapid knowledge evolution, while LLMs show limitations in retrieval and reasoning, making it difficult to uncover cross-document associations and reasoning pathways. To address these issues, we propose a pipeline that uses LLMs to construct a biomedical knowledge graph (BioStrataKG) from large-scale articles and builds a cross-document question-answering dataset (BioCDQA) to evaluate latent knowledge retrieval and multi-hop reasoning. We then introduce Integrated and Progressive Retrieval-Augmented Reasoning (IP-RAR) to enhance retrieval accuracy and knowledge reasoning. IP-RAR maximizes information recall through Integrated Reasoning-based Retrieval and refines knowledge via Progressive Reasoning-based Generation, using self-reflection to achieve deep thinking and precise contextual understanding. Experiments show that IP-RAR improves document retrieval F1 score by 20% and answer generation accuracy by 25% over existing methods. This framework helps doctors efficiently integrate treatment evidence for personalized medication plans and enables researchers to analyze advancements and research gaps, accelerating scientific discovery and decision-making.
Problem

Research questions and friction points this paper is trying to address.

Constructing biomedical knowledge graphs with complex terminology and data heterogeneity
Improving retrieval and reasoning capabilities of LLMs in biomedical contexts
Enhancing cross-document associations and multi-hop reasoning for clinical support
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLMs construct biomedical knowledge graph BioStrataKG
Integrated Progressive Retrieval-Augmented Reasoning IP-RAR
Self-reflection enhances retrieval and reasoning accuracy
🔎 Similar Papers
No similar papers found.
Yichun Feng
Yichun Feng
University of Chinese Academy of Sciences, Institute of Automation
Large Language ModelsKnowledge GraphsReinforcement LearningBioinformatics
J
Jiawei Wang
Department of EEIS, University of Science and Technology of China, Hefei 230026, China
R
Ruikun He
BYHEALTH Institute of Nutrition & Health, Guangzhou 510663, China
L
Lu Zhou
Guangzhou National Laboratory, No. 9 XingDaoHuanBei Road, Guangzhou International Bio Island, Guangzhou 510005, China
Yixue Li
Yixue Li
SIBS, CAS
Bioinformatics