Growing Hypergraphs with Homophily

📅 2026-07-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitation of existing hypergraph models that commonly assume conditional independence among hyperedges, which fails to capture the intricate dependencies between hyperedges and node labels as well as homophily mechanisms in real-world higher-order systems. We propose a growing dynamic hypergraph generative model that introduces inter-hyperedge dependencies through noisy copying of existing hyperedges and incorporates a multi-label-driven homophily mechanism to produce tunable assortative structures. By relaxing the hyperedge independence assumption, our model offers theoretical interpretability and supports likelihood-based parameter inference and community detection. We develop a stochastic EM algorithm for parameter estimation, integrate simulated annealing for community discovery, and analytically characterize the power-law degree distribution and the asymptotic behavior of intra-hyperedge label joint distributions. Experiments demonstrate that the model effectively captures homophily on both synthetic and real datasets, significantly outperforming edge-independent approaches in community detection.
📝 Abstract
There are many extant models of hypergraphs with interactions governed by attribute-based homophily between nodes, but most assume independence between edges conditional on node parameters. Relaxing this assumption, we study a mechanistic model of growing hypergraphs in which edge formation is influenced by both previous edges and binary node labels. Edges form in this model as noisy copies of previous edges, where the transmission of nodes from one edge to the next depends on multiple homophilic mechanisms between labels. These homophilic mechanisms give rise to tunable assortative structure in the hypergraph. We derive a power law for the degree distribution in this model and describe the long-term dynamics of the joint distribution of labels contained in edges. Our model defines a likelihood over a labeled hypergraph, allowing us to use standard maximum-likelihood techniques to structure algorithms. We estimate the model parameters on synthetic and real data via (stochastic) expectation maximization. These estimates give statistically-principled descriptions of the operation of homophily in empirical polyadic systems. We also demonstrate an approach to community detection via simulated annealing which, though computationally expensive, achieves competitive results on both synthetic data and certain empirical data sets known to be challenging to community detection techniques based on the edge-independence assumption. Our findings highlight the benefits of incorporating edge- and label-dependence in higher-order modeling and data analysis, and point to several directions for future work.
Problem

Research questions and friction points this paper is trying to address.

hypergraphs
homophily
edge dependence
node labels
assortative structure
Innovation

Methods, ideas, or system contributions that make the work stand out.

hypergraph growth
homophily
edge dependence
maximum likelihood estimation
community detection
🔎 Similar Papers
No similar papers found.
V
Violet Ross
Department of Computer Science, University of Colorado at Boulder, Boulder, CO, USA
F
Francis Cataldo
Department of Computer Science, Middlebury College, Middlebury, VT, USA
P
Philip S. Chodrow
Department of Computer Science, Middlebury College, Middlebury, VT, USA