CoJEPA: Combining Contrastive Learning and JEPA for Global-Local Music Representations
本文提出CoJEPA方法,结合对比学习和JEPA解决音乐表示中的全局-局部问题,通过共享骨干网络联合训练以获得更丰富的音乐表示。
本文提出CoJEPA方法,结合对比学习和JEPA解决音乐表示中的全局-局部问题,通过共享骨干网络联合训练以获得更丰富的音乐表示。
This study presents the first exploration into the detectability of AI-generated tracks within music co-created by humans and artificial intelligence. Addressing the challenge that general-purpose source separation methods often fail to reliably recover AI-specific artifacts, the authors propose a parallel detection architecture that operates without requiring full track separation. The approach leverages a neural audio codec to simulate the mixing process and combines short-time audio block analysis, relative energy estimation, and a binary classifier to directly identify AI-generated components within the mixed audio signal. Experiments on the MUSDB18-HQ dataset demonstrate promising track-level detection performance, confirming the feasibility of effectively discerning AI-generated content even in complex musical mixtures.
This work addresses the regulatory challenges posed by the proliferation of AI-generated music from unknown models by proposing an unsupervised, zero-shot detection framework that effectively distinguishes authentic from synthetic audio without relying on prior knowledge of generative models. The method integrates artifact-based feature extraction, non-negative matrix factorization (NMF), and a combination of one-class classification with unsupervised clustering strategies, thereby introducing zero-shot learning to AI music detection for the first time in a systematic manner. Experimental results demonstrate that the approach achieves strong performance in both binary authenticity discrimination and multi-class clustering of unseen AI-generated music sources, making it well-suited for monitoring high-purity synthetic content at scale within large music repositories.
This study addresses the lack of efficient, scalable, and human-aligned automatic evaluation methods for natural language responses in conversational music recommendation systems. It presents the first empirical investigation into the alignment between large language models employed as judges (LLM-as-a-Judge) and human experts, specifically along the dimensions of personalization and explanation quality. Through multi-turn dialogue sampling, generation of responses via four types of instruction-tuned models, expert human ratings, and bootstrap-based correlation analysis, the work demonstrates that LLM judges exhibit moderate positive correlation with human assessments—significantly outperforming conventional baselines. The findings further reveal the impact of model scale and contextual information on judging performance, offering practical guidance for model selection and deployment conditions in real-world applications.
This work addresses the overlooked issue of fairness disparities among sensitive attribute groups in multi-hop graph structures when promoting inter-group connections for link prediction. The paper introduces the concept of k-hop fairness and, for the first time, formalizes a structural bias metric that accounts for multi-hop distances, thereby revealing the intrinsic dependence between fairness and graph topology. This approach transcends the limitations of conventional methods that focus solely on first-order neighborhoods or pairwise fairness. By designing preprocessing and postprocessing strategies based on graph rewiring, the proposed method effectively mitigates multi-hop structural bias on standard benchmarks. Experimental results demonstrate that the approach significantly outperforms existing baselines in terms of multi-hop fairness while maintaining competitive link prediction performance.
本文提出CoJEPA方法,结合对比学习和JEPA解决音乐表示中的全局-局部问题,通过共享骨干网络联合训练以获得更丰富的音乐表示。
This study presents the first exploration into the detectability of AI-generated tracks within music co-created by humans and artificial intelligence. Addressing the challenge that general-purpose source separation methods often fail to reliably recover AI-specific artifacts, the authors propose a parallel detection architecture that operates without requiring full track separation. The approach leverages a neural audio codec to simulate the mixing process and combines short-time audio block analysis, relative energy estimation, and a binary classifier to directly identify AI-generated components within the mixed audio signal. Experiments on the MUSDB18-HQ dataset demonstrate promising track-level detection performance, confirming the feasibility of effectively discerning AI-generated content even in complex musical mixtures.
This work addresses the regulatory challenges posed by the proliferation of AI-generated music from unknown models by proposing an unsupervised, zero-shot detection framework that effectively distinguishes authentic from synthetic audio without relying on prior knowledge of generative models. The method integrates artifact-based feature extraction, non-negative matrix factorization (NMF), and a combination of one-class classification with unsupervised clustering strategies, thereby introducing zero-shot learning to AI music detection for the first time in a systematic manner. Experimental results demonstrate that the approach achieves strong performance in both binary authenticity discrimination and multi-class clustering of unseen AI-generated music sources, making it well-suited for monitoring high-purity synthetic content at scale within large music repositories.
This study addresses the lack of efficient, scalable, and human-aligned automatic evaluation methods for natural language responses in conversational music recommendation systems. It presents the first empirical investigation into the alignment between large language models employed as judges (LLM-as-a-Judge) and human experts, specifically along the dimensions of personalization and explanation quality. Through multi-turn dialogue sampling, generation of responses via four types of instruction-tuned models, expert human ratings, and bootstrap-based correlation analysis, the work demonstrates that LLM judges exhibit moderate positive correlation with human assessments—significantly outperforming conventional baselines. The findings further reveal the impact of model scale and contextual information on judging performance, offering practical guidance for model selection and deployment conditions in real-world applications.
This work addresses the overlooked issue of fairness disparities among sensitive attribute groups in multi-hop graph structures when promoting inter-group connections for link prediction. The paper introduces the concept of k-hop fairness and, for the first time, formalizes a structural bias metric that accounts for multi-hop distances, thereby revealing the intrinsic dependence between fairness and graph topology. This approach transcends the limitations of conventional methods that focus solely on first-order neighborhoods or pairwise fairness. By designing preprocessing and postprocessing strategies based on graph rewiring, the proposed method effectively mitigates multi-hop structural bias on standard benchmarks. Experimental results demonstrate that the approach significantly outperforms existing baselines in terms of multi-hop fairness while maintaining competitive link prediction performance.