UniVerse: Benchmarking and Enhancing LALMs on Culturally Inclusive Low-Resource Music Understanding

📅 2026-08-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决大型音频-语言模型在文化多样性音乐理解上的不足,研究通过构建包含5042个问答对的基准测试集UniVerseBench及自动化生成的训练数据集UniVerseSet,采用不平衡学习策略改进模型性能。
📝 Abstract
Recent advances in large audio-language models (LALMs) have significantly improved performance in tasks such as music captioning, genre classification, and sound event detection. However, limited attention has been paid to improving their adaptability across diverse musical traditions, particularly folk music rooted in distinct cultural contexts. Folk-music traditions are typically resource-scarce, unevenly represented across regions, and poorly documented. Even when such samples appear in large-scale pre-training, LALMs often fail to capture their structural and stylistic characteristics, partly due to the absence of dedicated evaluation protocols and training solutions. To address these limitations, we introduce UniVerse, a reproducible solution for low-resource music understanding. Specifically, we propose UniVerseBench, a benchmark of 5,042 Q&A pairs across more than 38 cultural and linguistic entities, constructed via an expert-guided yet highly automated pipeline. In parallel, we construct a fully automated, model-generated multi-turn dialogue training dataset UniVerseSet. By training LALMs on UniVerseSet, we systematically adapt and investigate representative multimodal imbalance learning strategies across both dense and Mixture-of-Experts (MoE) architectures. Experimental results indicate that fully automated data curation combined with imbalance-aware training yields non-trivial improvements, but models still struggle to capture fine-grained acoustic features, indicating a gap between surface-level alignment and deep musical comprehension.
Problem

Research questions and friction points this paper is trying to address.

LALMs
Culturally Inclusive
Low-Resource Music Understanding
Folk Music
Structural and Stylistic Characteristics
Innovation

Methods, ideas, or system contributions that make the work stand out.

LALMs
UniVerseBench
UniVerseSet
imbalance learning strategies
cultural inclusiveness
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Ziya Zhou
Ziya Zhou
The Hong Kong University of Science and Technology
Music TechnologyNatural Language Processing
Shangda Wu
Shangda Wu
Tencent
Symbolic Music GenerationMusic Information RetrievalMultimodal Learning
S
Shenyang Xu
Independent Researcher
Y
Yutong Zheng
Independent Researcher
D
Dafang Liang
Independent Researcher
S
Suin Chung
Sogang University, Seoul, Korea
Danbinaerin Han
Danbinaerin Han
KAIST
MIR(Music Information Retrieval)Korean traditional music
J
Junyan Jiang
NYU Shanghai, China
Yongyi Zang
Yongyi Zang
Smule, Inc.
Computer AuditionSpeech ProcessingMusic Information RetrievalMusic Composition
Ruibin Yuan
Ruibin Yuan
HKUST
Artificial IntelligenceMusic GenerationMusic Information RetrievalComputer Music
R
Rongxiu Zhong
JIUTIAN Research, China Mobile, Beijing, China; The State Key Laboratory of Multimedia Information Processing, Peking University, Beijing, China
S
Shilei Zhang
JIUTIAN Research, China Mobile, Beijing, China; The State Key Laboratory of Multimedia Information Processing, Peking University, Beijing, China
Junlan Feng
Junlan Feng
Chief Scientist at China Mobile Research
Natural LanguageMachine LearningSpeech ProcessingData Mining
J
Jinglei Liu
China Mobile (Hong Kong) Innovation Research Institute, Hong Kong SAR, China
Haotian Zhou
Haotian Zhou
MLSys@ByteDance Seed
Machine Learning SystemLLMMachine Learning
Z
Zijin Li
Central Conservatory of Music, Beijing, China
Dasaem Jeong
Dasaem Jeong
Sogang University
Music Information RetrievalExpressive Performance ModelingMachine Learning
Wei Xue
Wei Xue
HKUST
Audio ProcessingAI MusicFoundation ModelsGenerative AIMultimodal
Y
Yike Guo
HKUST, Hong Kong SAR, China