TAHB: A Comprehensive Benchmark for Text-Attributed Hypergraph Learning

📅 2026-08-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of public benchmarks integrating language models with hypergraph learning by introducing TAHB, the first text-attributed hypergraph benchmark. Comprising ten real-world datasets, TAHB supports dual evaluation paradigms encompassing both LLM augmentation and prediction tasks. Systematic experiments validate the efficacy of text-aware hypergraph representation learning, demonstrating that LLM-enhanced semantics significantly improve model performance and that joint structural-textual modeling constitutes the optimal predictive strategy. By filling a critical gap in the literature, TAHB provides a standardized evaluation platform and essential empirical evidence for the deep integration of hypergraph learning and large language models, thereby facilitating future research at this emerging intersection.
📝 Abstract
Hypergraphs effectively model higher-order groupwise relationships beyond pairwise interactions, while pretrained language models (PLMs) and large language models (LLMs) provide rich semantic understanding from textual attributes. However, research on combining language models with hypergraph learning remains limited due to the lack of public text-attributed hypergraph benchmarks. To address this limitation, we present TAHB (Text-Attributed Hypergraph Benchmark), the first public benchmark integrating hypergraph structures and raw textual attributes. TAHB contains 10 real-world datasets from four domains - e-commerce, academia, movies, and politics networks - enabling systematic evaluation of text-aware hypergraph representation learning. Experimental results show that TAHB preserves key structural properties of real-world hypergraphs and consistently reproduces performance tendencies observed in existing benchmarks. Furthermore, experiments under both LLM-as-Enhancer and LLM-as-Predictor settings demonstrate that LLM-enhanced textual semantics improve hypergraph learning performance, while structural and textual information jointly provide the best setting for LLM-based prediction. Our benchmark provides a foundation for future research at the intersection of hypergraph learning and language models.
Problem

Research questions and friction points this paper is trying to address.

Text-Attributed Hypergraph
Benchmark
Hypergraph Learning
Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Text-Attributed Hypergraph
Benchmark
Large Language Models
Hypergraph Learning
LLM-as-Enhancer
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
D
David Yoon Suk Kang
Chungbuk National University, South Korea
J
JungHyun Kim
Hanyang University, South Korea
J
Juhyun Jeon
Hanyang University, South Korea
Sang-Wook Kim
Sang-Wook Kim
Professor of Computer Science and Engineering, Hanyang University
DatabasesData MiningSocial Network AnalysisRecommendation