From Queries to Narratives: Cultural Heritage Data Stories for Knowledge Graph Exploration and Quality Assessment
本文提出数据故事方法,通过结合解释性文本、图像与可执行SPARQL查询及其可视化结果,降低文化遗产知识图谱的探索门槛,并促进数据质量评估。
本文提出数据故事方法,通过结合解释性文本、图像与可执行SPARQL查询及其可视化结果,降低文化遗产知识图谱的探索门槛,并促进数据质量评估。
该研究构建了zbMATH开放知识图谱,通过整合专家策划的语义内容来追踪超过250年的数学研究成果,支持对数学概念和学术关系的细粒度探索。
This study addresses the unclear community structure within the mathematical software ecosystem and the challenge of predicting paper-software associations. We construct a software co-occurrence network to reveal community heterogeneity and formulate association prediction as a multi-label classification task. Comparative experiments integrating Mathematics Subject Classification (MSC) metadata with title embeddings demonstrate that structured MSC features significantly outperform purely semantic representations in precision-recall trade-offs. This work not only elucidates the topological structure of the mathematical software landscape but also validates the critical value of domain-specific metadata in academic recommendation systems, thereby providing effective support for software discovery.
This study addresses the challenges of semantic representation and reuse of the Central Federal Card Index within post-war German reparations archives. To tackle this, the authors propose a modular two-layer ontology architecture that decouples domain semantics—grounded in the Basic Formal Ontology (BFO) realist framework—from interoperability structures aligned with established standards such as RiC-O, PROV-O, and PiCo. This design ensures logical rigor while enabling effective semantic integration and knowledge graph construction in digital humanities infrastructures. The resulting BZKO ontology not only facilitates robust semantic modeling of historical archival data but also establishes a foundation for incorporating spatiotemporal uncertainty in future extensions.
Scientific formulas inherently encode both syntactic structure and semantic meaning, yet these two aspects exhibit significant misalignment in their native representation spaces, limiting the performance of cross-modal retrieval. This work is the first to systematically uncover the weakly observable correspondence between formula syntax and semantics and proposes an explicit alignment approach. Specifically, it employs a graph neural network to encode syntactic structures and a text encoder to model semantic content, integrating them into a unified representation space through contrastive learning. The resulting aligned representations effectively bridge the modality gap and substantially enhance cross-modal retrieval performance, thereby demonstrating the critical role of explicit representation learning in mathematical formula understanding.
本文提出数据故事方法,通过结合解释性文本、图像与可执行SPARQL查询及其可视化结果,降低文化遗产知识图谱的探索门槛,并促进数据质量评估。
该研究构建了zbMATH开放知识图谱,通过整合专家策划的语义内容来追踪超过250年的数学研究成果,支持对数学概念和学术关系的细粒度探索。
This study addresses the unclear community structure within the mathematical software ecosystem and the challenge of predicting paper-software associations. We construct a software co-occurrence network to reveal community heterogeneity and formulate association prediction as a multi-label classification task. Comparative experiments integrating Mathematics Subject Classification (MSC) metadata with title embeddings demonstrate that structured MSC features significantly outperform purely semantic representations in precision-recall trade-offs. This work not only elucidates the topological structure of the mathematical software landscape but also validates the critical value of domain-specific metadata in academic recommendation systems, thereby providing effective support for software discovery.
This study addresses the challenges of semantic representation and reuse of the Central Federal Card Index within post-war German reparations archives. To tackle this, the authors propose a modular two-layer ontology architecture that decouples domain semantics—grounded in the Basic Formal Ontology (BFO) realist framework—from interoperability structures aligned with established standards such as RiC-O, PROV-O, and PiCo. This design ensures logical rigor while enabling effective semantic integration and knowledge graph construction in digital humanities infrastructures. The resulting BZKO ontology not only facilitates robust semantic modeling of historical archival data but also establishes a foundation for incorporating spatiotemporal uncertainty in future extensions.
Scientific formulas inherently encode both syntactic structure and semantic meaning, yet these two aspects exhibit significant misalignment in their native representation spaces, limiting the performance of cross-modal retrieval. This work is the first to systematically uncover the weakly observable correspondence between formula syntax and semantics and proposes an explicit alignment approach. Specifically, it employs a graph neural network to encode syntactic structures and a text encoder to model semantic content, integrating them into a unified representation space through contrastive learning. The resulting aligned representations effectively bridge the modality gap and substantially enhance cross-modal retrieval performance, thereby demonstrating the critical role of explicit representation learning in mathematical formula understanding.