🤖 AI Summary
This paper addresses the challenge of modeling degree sequences in multidisciplinary citation networks, where community heterogeneity—such as disparities in growth rates, reference list lengths, and preferential citation tendencies—complicates traditional modeling. We propose an extended 3DSI model that integrates heterogeneous growth, preferential attachment, and continuous-time Markov processes, augmented with extreme-value statistics theory. For the first time, we analytically prove that within-community citation distributions converge to the Pareto Type II distribution. This derivation yields interpretable, closed-form metrics for inequality (Gini coefficient) and preference, along with a principled parameter estimation framework. Empirical validation on real-world multidisciplinary citation data demonstrates that our model significantly outperforms the classical Price model, achieving high-fidelity cross-disciplinary degree distribution fitting, quantitative comparison of citation inequality across fields, and mechanistic attribution of underlying drivers.
📝 Abstract
We introduce a new analytical framework for modelling degree sequences in individual communities of real-world networks, e.g., citations to papers in different fields. Our work is inspired by Price's model and its recent generalisation called 3DSI (three dimensions of scientific impact), which assumes that citations are gained partly accidentally, and to some extent preferentially. Our generalisation is motivated by existing research indicating significant differences between how various scientific disciplines grow, namely, minding different growth ratios, average reference list lengths, and preferential citing tendencies. Extending the 3DSI model to heterogeneous networks with a community structure allows us to devise new analytical formulas for, e.g., citation number inequality and preferentiality measures. We show that the distribution of citations in a community tends to a Pareto type II distribution. We also present analytical formulas for estimating its parameters and Gini's index. The new model is validated on real citation networks.