🤖 AI Summary
研究通过重组和再利用现有数据集解决AI研究中高质量数据需求与易获取数据源枯竭的问题,发现数据再利用与更高的科学颠覆性和引用影响相关。
📝 Abstract
Technological advancements are enabling increasingly systematic and large-scale data collection across all areas of science, driving scientific innovation. In particular, AI research exemplifies this trend, having advanced rapidly through the assembly of massive datasets used to train and evaluate machine learning models. However, the escalating demand for data, the difficulty of creating high-quality datasets, and the exhaustion of easily accessible data sources in AI research raise important questions about how to maximize the value of existing datasets through recombination and repurposing. Here, we draw on two theoretical frameworks---recombinational novelty and transformational creativity---to examine the practice of data repurposing and its scientific impact. Focusing on AI, we analyze scientific outcomes associated with data repurposing across more than 10,000 machine learning papers. First, we find that although most repurposed datasets do not achieve broad visibility in the short term, data repurposing is associated with greater disruption. Second, when repurposed data is adopted by subsequent research, the repurposing paper is associated with higher disruption and increased citation impact. Third, repurposing teams tend to be more experienced, more institutionally prestigious, and involve academic--industry collaboration. However, team characteristics poorly predict which repurposed datasets will be adopted by the community. These findings suggest that data repurposing may be an important approach to scientific discovery, and that its successful adoption is more common among larger teams and collaborations spanning academia and industry.