WhichTok? Comparing Three TikTok Data Acquisition Tools

📅 2026-08-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the reproducibility crisis in TikTok research stemming from the platform’s opaque algorithms and the absence of standardized data collection methods. It presents the first systematic comparison of three widely used tools—TikTok’s official Research API, Pyktok, and Apify—evaluating their performance across five data endpoints: users, hashtags, keywords, comments, and related videos. Findings reveal consistent results only for user-level data, with substantial discrepancies observed across all other dimensions. Critically, none of the tools enable truly random sampling, thereby introducing potentially significant but unquantifiable biases. The work underscores fundamental limitations in the representativeness and consistency of current data collection approaches and offers actionable recommendations to enhance research transparency and ethical rigor in future studies of TikTok and similar algorithmically mediated platforms.
📝 Abstract
TikTok's global growth has made it a prime platform for both entertainment and political discourse, prompting increased social science research. However, this rapidly evolving research field faces a fundamental reproducibility crisis. TikTok's opaque algorithmic systems hinder researchers from drawing meaningful empirical inferences, while the lack of standardized data collection methods compounds these challenges. This study addresses these methodological gaps by systematically comparing three data collection tools - the official TikTok Research API, Pyktok, and Apify. We evaluated five endpoints: User, Hashtag, Keyword, Comment, and Related Video. Results show substantial cross-tool differences, especially for hashtag and keyword searches. The Research API uses back-end API calls, whereas Apify and Pyktok rely on front-end web scraping, producing systematic differences in the time periods and popularity levels represented in retrieved content. The three tools yielded comprehensive and consistent results only for the user endpoint. Our results question whether these tools can acquire truly random[-ized] samples, as they introduce methodological confounds that may compromise research validity in ways not yet fully understood. Based on these results, we offer methodological, transparency, and ethical recommendations and guidelines to increase TikTok research quality.
Problem

Research questions and friction points this paper is trying to address.

reproducibility crisis
data acquisition
TikTok research
algorithmic opacity
methodological confounds
Innovation

Methods, ideas, or system contributions that make the work stand out.

TikTok data collection
Research API
Web scraping
Reproducibility
Methodological bias
🔎 Similar Papers