HERMES: a multi-agent framework for structured knowledge extraction from ultra-long documents in geoscience

๐Ÿ“… 2026-08-14
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the challenges posed by unstructured content and complex layouts in lengthy geoscience literature that impede knowledge computing. We propose a scalable multi-agent framework integrating text and chart parsing to establish a unified document-level extraction pipeline incorporating domain constraints, validation rules, and evidence tracing for cross-domain transferability. Experiments demonstrate that the system extracted over 30,000 entities and 450,000 attributes with an F1 score of 0.90, achieving a sixfold efficiency improvement. Furthermore, we released a cross-domain validated database, effectively resolving the difficulties associated with structured knowledge extraction from long-form scientific documents.
๐Ÿ“ Abstract
Authoritative scientific knowledge in geoscience remains largely trapped in legacy monographs and historical literature, where unstructured text and complex layouts hinder computational access. We introduce HERMES, a scalable multi-agent framework that extracts structured data from ultra-long scientific documents. Using a coordinating large language model, HERMES integrates domain constraints, validation rules and evidence tracing within a unified document-level extraction process that incorporates parsed text, tables, figures and captions. Applied to the 55-volume Treatise on Invertebrate Paleontology, the system produced a structured database of 32,277 fossil taxonomic entities and 451,878 attributes, released online at https://treatise.geolex.org. Extraction performance remained stable across fossil groups (average F1 scores of approximately 0.90 for entities and 0.91 for attributes), improving per-volume efficiency approximately sixfold relative to the tested fully manual baseline. Evaluation in palaeomagnetism and geochemistry, conducted without additional model training, demonstrated transfer across distinct geoscience domains. This work provides a practical pathway to transform historical scientific literature into FAIR-oriented structured data, offering a sustainable infrastructure for data-intensive disciplines and large-scale knowledge integration.
Problem

Research questions and friction points this paper is trying to address.

structured knowledge extraction
ultra-long documents
geoscience
unstructured text
FAIR data
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-agent framework
Structured knowledge extraction
Ultra-long documents
Evidence tracing
Cross-domain transfer
๐Ÿ”Ž Similar Papers
๐Ÿ’ผ Related Jobs
No related jobs found.
Z
Ziqi Song
Research Center for Computational Earth and Space Science, Zhejiang Laboratory, Hangzhou, China
Z
Zongyuan Xiang
Research Center for Computational Earth and Space Science, Zhejiang Laboratory, Hangzhou, China
J
James G. Ogg
School of Electrical & Computer Engineering, Purdue University, Indiana, USA; Zhejiang Deep-time Digital Earth International Research Center, Hangzhou, China
B
Bruce S. Lieberman
Biodiversity Institute, University of Kansas, USA
G
Gabi Ogg
Geologic TimeScale Foundation, Indiana, USA
N
Natalia Lรณpez Carranza
Biodiversity Institute, University of Kansas, USA
W
Wen Du
Department of Earth and Environmental Sciences, University of Illinois Chicago, Chicago, IL, USA
Yufei Ye
Yufei Ye
Stanford University
Computer Vision
S
Shuan Li
Research Center for Computational Earth and Space Science, Zhejiang Laboratory, Hangzhou, China
Z
Zhong Peng
Research Center for Computational Earth and Space Science, Zhejiang Laboratory, Hangzhou, China
Shaoqi Yu
Shaoqi Yu
Zhejiang University
machine learninggenerative model
J
Juye Wei
Research Center for Computational Earth and Space Science, Zhejiang Laboratory, Hangzhou, China
Y
Ying Zhou
Research Center for Computational Earth and Space Science, Zhejiang Laboratory, Hangzhou, China
J
Jieping Ye
Research Center for Computational Earth and Space Science, Zhejiang Laboratory, Hangzhou, China
Jiang Yang
Jiang Yang
Southern University of Science and Technology
numerical analysisscientific computing