Building evidence-based knowledge graphs from full-text literature for disease-specific biomedical reasoning
Existing biomedical knowledge resources are often limited to unstructured text or flat triples that omit critical contextual details such as study design, source provenance, and quantitative support, thereby hindering fine-grained reasoning. This work proposes a large language model–assisted end-to-end framework that extracts experimental findings from full-text literature as structured evidence nodes, standardizes biomedical entities, evaluates evidence quality, and constructs disease-specific knowledge graphs via typed semantic relations. To our knowledge, this is the first approach to systematically build structured evidence graphs from full-text articles while preserving study design, source attribution, and quantitative backing. The authors release two high-quality datasets, EvidenceNet-HCC and EvidenceNet-CRC, with component-level validation accuracy of 98.3%, demonstrating significant performance gains in retrieval-augmented question answering, link prediction, and target prioritization tasks.