MUUNRiver-Bench: Diagnosing Relation-Dependent Music Retrieval with Multimodal Instructions
研究针对音乐检索中关系依赖性问题,通过MUUNRiver-Bench基准,使用多模态指令来定义相关性,揭示了不同模型在处理特定任务时的偏好差异。
研究针对音乐检索中关系依赖性问题,通过MUUNRiver-Bench基准,使用多模态指令来定义相关性,揭示了不同模型在处理特定任务时的偏好差异。
为解决高质量音频到乐谱转录数据稀缺问题,研究通过完全合成的数据和Transformer模型预训练方法,提高跨乐器转录性能。
该研究通过引入Dissonance Spectrum来显式建模同时频率成分间的关系,以改进音乐理解,并在多项音乐理论测试中表现出色。
为解决大型音频-语言模型在文化多样性音乐理解上的不足,研究通过构建包含5042个问答对的基准测试集UniVerseBench及自动化生成的训练数据集UniVerseSet,采用不平衡学习策略改进模型性能。
This study addresses the rapid yet uneven growth of artificial intelligence research in music, where attention across tasks lacks systematic measurement. To characterize this imbalance, the authors construct a joint taxonomy encompassing 12 application domains and 11 technical methodologies, analyzing 6,839 publications from 2015 to 2026. They propose a novel four-dimensional profile of research attention—comprising technical investment, method allocation, methodological diversity, and adoption lag of cutting-edge techniques—and apply bibliometric analysis combined with time-lag modeling. The findings reveal that generative tasks adopt state-of-the-art methods rapidly (average lag: 0.33 years), whereas applications in education, health, and governance exhibit substantial delays (4.33 and 5.00 years, respectively), highlighting significant disparities in research resource allocation and filling a critical gap in assessing equity within AI-driven music research.
研究针对音乐检索中关系依赖性问题,通过MUUNRiver-Bench基准,使用多模态指令来定义相关性,揭示了不同模型在处理特定任务时的偏好差异。
为解决高质量音频到乐谱转录数据稀缺问题,研究通过完全合成的数据和Transformer模型预训练方法,提高跨乐器转录性能。
该研究通过引入Dissonance Spectrum来显式建模同时频率成分间的关系,以改进音乐理解,并在多项音乐理论测试中表现出色。
为解决大型音频-语言模型在文化多样性音乐理解上的不足,研究通过构建包含5042个问答对的基准测试集UniVerseBench及自动化生成的训练数据集UniVerseSet,采用不平衡学习策略改进模型性能。
This study addresses the rapid yet uneven growth of artificial intelligence research in music, where attention across tasks lacks systematic measurement. To characterize this imbalance, the authors construct a joint taxonomy encompassing 12 application domains and 11 technical methodologies, analyzing 6,839 publications from 2015 to 2026. They propose a novel four-dimensional profile of research attention—comprising technical investment, method allocation, methodological diversity, and adoption lag of cutting-edge techniques—and apply bibliometric analysis combined with time-lag modeling. The findings reveal that generative tasks adopt state-of-the-art methods rapidly (average lag: 0.33 years), whereas applications in education, health, and governance exhibit substantial delays (4.33 and 5.00 years, respectively), highlighting significant disparities in research resource allocation and filling a critical gap in assessing equity within AI-driven music research.