FUBU-EPSTEIN: A Large-Scale Twitter Dataset on the Jeffrey Epstein Case and Its Global Public Discourse (2019-2023)

📅 2026-08-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of large-scale, structured social media datasets enabling computational analysis of global public discourse surrounding the Jeffrey Epstein case. We construct a multidimensional Twitter corpus spanning 2019 to 2023, comprising 54.38 million deduplicated tweets and associated user interaction graphs, thereby facilitating the first longitudinal tracking of sustained public discussion on this topic. Leveraging the Qwen2.5-7B-Instruct large language model, we automatically annotate key dimensions—including sentiment, conspiracy-theory stance, and toxicity—while implementing rigorous data anonymization and de-identification protocols to ensure privacy compliance. The project releases a publicly available dataset containing 46.15 million internal relationships and offers controlled access to raw data for collaborative research, balancing scholarly innovation with ethical standards.
📝 Abstract
The criminal case of Jeffrey Epstein has generated a complex, long-running global discourse on digital platforms, characterized by punctuated attention shocks, conspiracy theories, and blame attribution. To facilitate the computational study of these dynamics, we introduce the FUBU-EPSTEIN dataset, a large-scale, multi-dimensional research corpus of 54.38 million Twitter statuses authored by 7.11 million users and collected between August 2019 and April 2023. The source corpus was captured continuously in near-real time and enriched with a directed social contact graph of 37.03 million edge rows and annotations from Qwen2.5-7B-Instruct covering sentiment, conspiracy and misinformation stance, toxicity, moral emotion, and related dimensions. For public distribution, we created a textless, de-identified derivative that retains one row for every deduplicated status, categorical and numeric annotations, coarse temporal information, and 46.15 million internal status relationships. It excludes tweet text, original post and user identifiers, handles, profile fields, exact timestamps, and reverse mappings. The resulting release supports longitudinal content and diffusion analyses while reducing disclosure and platform-content redistribution risks. To request access to raw data for collaborative research under ethical and legal safeguards, contact fubu.dataset@gmail.com.
Problem

Research questions and friction points this paper is trying to address.

Jeffrey Epstein
public discourse
conspiracy theories
social media dataset
misinformation
Innovation

Methods, ideas, or system contributions that make the work stand out.

large-scale dataset
social contact graph
LLM-based annotation
privacy-preserving release
longitudinal discourse analysis
🔎 Similar Papers
No similar papers found.
M
Michael Kreil
Published as Independent Researcher
T
Tristan Manfred Stöber
Published as Independent Researcher
Daniel Thilo Schroeder
Daniel Thilo Schroeder
SINTEF, Oslo Metropolitan University
Computational Social ScienceMisinformation