๐ค AI Summary
This study addresses the challenge of trustworthy scientific data discovery in the era of data explosion by introducing the concept of โdata socio-technical lifeโ and constructing a data usage graph that dynamically links datasets to associated publications, researchers, institutions, topics, software, and models within scientific practice. By integrating scholarly knowledge graph construction, context-aware trustworthiness assessment, and national data platform interoperability techniques, the project implements a prototype โData Insight Discovery Systemโ within the National Data Platform (NDP). The system demonstrates the effectiveness of this graph-based approach as a novel research infrastructure for enhancing the evaluation of data suitability and credibility, thereby overcoming the limitations inherent in traditional metadata frameworks.
๐ Abstract
Artificial intelligence is changing the scale and tempo of scientific inquiry. Models can now search, integrate, and reason over data far beyond data repositories familiar to any individual researcher. Yet this expansion creates a prior problem: before a model can produce a trustworthy scientific result, it must locate data that are appropriate for the question, sufficiently reliable for the intended analysis, and accompanied by enough context to support responsible interpretation. As data becomes increasingly abundant, the challenge of finding data has been overcome by the challenge of finding data that you can trust.
This paper explores how the social and empirical evidence that accumulates when data are used in research can be used, analogous to social trust networks, to determine fit for purpose and trust. Specifically, the paper explores data-usage graphs as a new layer of scientific data infrastructure. A data-usage graph connects datasets to the publications, people, institutions, topics, software, models, workflows, and other datasets through which they are produced and used. These connections reveal the {\it social life of data:} who has relied on a source, for which questions, in what combinations, with which methods, and with what observable impact. They can turn scattered traces of practice into data-usage descriptors that complement conventional metadata and support judgments of trust and fitness for purpose. The central claim is not that popularity establishes trust, but that this can be grown with appropriate contextual history. Usage evidence must therefore be combined with production quality, provenance, governance, semantic clarity, and community validation. The feasibility and value of data usage graphs is demonstrated by implementing the prototype data insights discovery service within the National Data Platform (NDP).