Extracting ontology-compliant knowledge from scientific text describing irradiated materials using large language models

📅 2026-09-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究使用大型语言模型从描述辐照材料的科学文本中提取符合本体的知识,解决了数据重用和结构化问题。
📝 Abstract
The quest for new materials increasingly relies on predictive models and comprehensive simulations that span scales from atomic to macroscopic levels. However, essential data necessary for these models and simulations are often embedded in scientific literature as unstructured text, limiting reusability and posing challenges for researchers seeking to leverage existing knowledge effectively. While extracting structured data from unstructured text using large language models is gaining popularity, traditional methods typically generate key-value pairs data with straightforward schemas. In contrast, we introduce eolas, a modular pipeline that uses large language models to automatically transform scientific documents into knowledge graphs aligned with a specified ontology. We demonstrate eolas effectiveness in extracting useful information for scientists studying materials designed to endure the extreme temperatures and radiation levels found in fusion reactors. While a human expert might spend between thirty to ninety minutes extracting relevant data from an article, eolas can generate high-quality knowledge graphs in just a few minutes. These are presented in a tabular format with faceted navigation for easy human validation. Additionally, we introduce the first benchmark dataset designed to assess large language models capabilities in constructing knowledge graphs within the domain of irradiated materials. The analysis of 168 experiments using our dataset, various large language models and prompting techniques provides key insights that we summarize into practical guidelines for effectively extracting knowledge graphs aligned with an input ontology.
Problem

Research questions and friction points this paper is trying to address.

unstructured text
predictive models
comprehensive simulations
data reusability
knowledge extraction
Innovation

Methods, ideas, or system contributions that make the work stand out.

large language models
knowledge graphs
ontology alignment
unstructured text extraction
benchmark dataset
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Marco Luca Sbodio
Marco Luca Sbodio
IBM Research
M
Marcos Martínez Galindo
IBM Research
Vanessa Lopez
Vanessa Lopez
Knowledge Media Institute, IBM Research Europe
Semantic WebQuestion AnsweringCity Data
B
Blanca Biel
Dpto. Física Atómica, Molecular y Nuclear, Instituto Carlos I de Física Teórica y Computacional, Univ. of Granada, Spain
P
Pablo Canca
Dpto. Física Atómica, Molecular y Nuclear, Univ. of Granada, Spain
P
Pedro Delgado
IFMIF DONES, Spain
J
Jesús I. Mendieta-Moreno
Instituto de Ciencia de Materiales de Madrid (ICMM), CSIC, Spain
R
Raphael Tack
IBM Research
M
Maria J. Caturla
Dpto. de Física, Facultad de Ciencias, Universidad de Alicante, Spain