Generating Synthetic Citation Networks with Communities

๐Ÿ“… 2026-04-28
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Existing synthetic graph generation methods struggle to simultaneously capture both the community structure and the near-acyclic nature of real citation networks, limiting the evaluation of community detection and network mining algorithms. This work systematically evaluates twelve directed graph generators on seven real-world citation networks and introduces an edge direction reversal strategy to break cycles and emulate citation flow. It further distinguishes, for the first time, between endogenous and exogenous mesoscale structural similarities, revealing that high-parameter models often overfit by memorizing community statistics. Building on these insights, we propose Citation Seederโ€”a novel algorithm combining a Price-Pareto iterative mechanism with a degree-corrected stochastic block model. Evaluated using only 26 multidimensional metrics, Citation Seeder achieves interpretable and efficient citation network generation with four orders of magnitude fewer parameters, preserving structural fidelity while significantly enhancing computational efficiency and generalization capability.
๐Ÿ“ Abstract
Generating realistic synthetic citation, patent, or component dependency networks is essential for benchmarking community detection, graph visualisation, and network data mining algorithms. We present the first systematic comparison of generators of directed graphs that are nearly acyclic and have a ground-truth community structure. We evaluate 12 methods across 7 real citation networks and 26 metrics. We propose the practice of reversing directions of edges in static generators to break cycles and induce a citation-like flow, which significantly improves the performance of a degree-corrected Stochastic Block Model. Our novel methodological approach to evaluating community detection benchmarks distinguishes between endogenous and exogenous mesoscopic similarities, with the latter proving more important. This distinction reveals that high-parameter models suffer from overfitting by memorising planted community statistics which lead to their failing to produce realistic networks. Finally, we introduce the Citation Seeder (CS) algorithm, an iterative generator grounded in the Price-Pareto model of citation networks, with interpretable parameters and O(N+E) runtime. CS achieves competitive results against the best-performing baselines while using up to four orders of magnitude fewer parameters and providing a clean framework for explaining and predicting a network's future growth.
Problem

Research questions and friction points this paper is trying to address.

synthetic citation networks
community detection
graph generation
benchmarking
directed acyclic graphs
Innovation

Methods, ideas, or system contributions that make the work stand out.

synthetic citation networks
community detection benchmarking
Citation Seeder algorithm
degree-corrected Stochastic Block Model
mesoscopic similarity
๐Ÿ”Ž Similar Papers
๐Ÿ’ผ Related Jobs
No related jobs found.
ล
ลukasz Brzozowski
Warsaw University of Technology, Faculty of Mathematics and Information Science, ul. Koszykowa 75, 00-662 Warsaw, Poland
Marek Gagolewski
Marek Gagolewski
Warsaw University of Technology
Machine LearningComputational StatisticsStatistical SoftwareData Fusion and Aggregation
G
Grzegorz Siudem
Warsaw University of Technology, Faculty of Physics, ul. Koszykowa 75, 00-662 Warsaw, Poland