🤖 AI Summary
This study addresses the challenges of missing data and insufficient analytical precision in academic publishing Article Processing Charge (APC) research. Leveraging web scraping and Wayback Machine historical snapshots, we constructed a comprehensive APC dataset spanning 2019–2025 across 14 publishers. Comprising nearly 70,000 records for 12,540 journals, this dataset provides longitudinal price lists in multiple currencies alongside journal metadata. By effectively bridging the gap in longitudinal APC data, this work significantly enhances the accuracy of cost estimations. Consequently, it offers high-quality empirical support for open access cost analysis and library collection development, establishing a robust foundation for future bibliometric and economic assessments of scholarly publishing.
📝 Abstract
This paper introduces a dataset of APCs produced from the price lists of 14 large scholarly publishers between 2019 and 2025. APC price lists were downloaded from publisher websites each year as well as via Wayback Machine snapshots to retrieve fees per journal per year. The dataset includes journal metadata, APC collection method, and annual APC price list information in several currencies (USD, EUR, GBP, CHF, JPY, CAD, AUD) for 12,540 unique journals and 69,856 journal-year combinations. The dataset was generated to allow for more precise analysis of APCs and can support library collection development and scientometric analysis estimating APCs paid in gold and hybrid OA journals.