🤖 AI Summary
This work addresses the challenge posed by unstructured textual descriptions of density functional theory (DFT) workflows in computational materials science, which hinder reproducibility and systematic comparison. We propose the first framework that integrates domain ontologies with large language models (LLMs) to automatically extract stacking fault energy computation protocols from scientific literature. Through a multi-stage text filtering pipeline and tailored prompt engineering, our approach aligns extracted workflows with established materials ontologies—including CMSO, ASMO, and PLDO—and constructs a structured atomRDF knowledge graph. This method achieves, for the first time, semantic structuring and cross-study alignment of DFT workflows, substantially enhancing the machine-readability, transparency, and reusability of computational materials data.
📝 Abstract
Reproducibility of computational results remains a challenge in materials science, as simulation workflows and parameters are often reported only in unstructured text and tables. While literature data are valuable for validation and reuse, the lack of machine-readable workflow descriptions prevents large-scale curation and systematic comparison. Existing text-mining approaches are insufficient to extract complete computational workflows with their associated parameters. An ontology-driven, large language model (LLM)-assisted framework is introduced for the automated extraction and structuring of computational workflows from the literature. The approach focuses on density functional theory-based stacking fault energy (SFE) calculations in hexagonal close-packed magnesium and its binary alloys, and uses a multi-stage filtering strategy together with prompt-engineered LLM extraction applied to method sections and tables. Extracted information is unified into a canonical schema and aligned with established materials ontologies (CMSO, ASMO, and PLDO), enabling the construction of a knowledge graph using atomRDF. The resulting knowledge graph enables systematic comparison of reported SFE values and supports the structured reuse of computational protocols. While full computational reproducibility is still constrained by missing or implicit metadata, the framework provides a foundation for organizing and contextualizing published results in a semantically interoperable form, thereby improving transparency and reusability of computational materials data.