DataParasite Enables Scalable and Repurposable Online Data Curation
This work proposes an open-source, modular data collection pipeline to address the labor-intensive, costly, and poorly reproducible nature of dataset construction from heterogeneous online sources in computational social science. The system employs lightweight natural language instructions to configure workflows, decomposing tabular data curation into entity-level search and structured extraction tasks. By leveraging large language model–based intelligent agents, it achieves a task-agnostic, reusable architecture that operates effectively even without predefined entity lists. Evaluated across multiple representative tasks, the approach attains high accuracy while reducing data acquisition costs by an order of magnitude compared to manual methods, substantially lowering both technical and human resource barriers to entry.