Lead Vocal Separation from Vocal Ensemble Mixtures Using Phoneme Alignment

📅 2026-09-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出了一种利用音素对齐信息辅助的模型来解决无伴奏合唱中主唱分离的问题,该方法基于改进的BS-RoFormer架构,并通过实验验证了其有效性。
📝 Abstract
Contemporary a cappella singing often has a lead-and-accompaniment texture, where the lead vocal (Vo) part carries the main melody and the remaining vocal parts provide accompaniment. Owing to their distinct roles, separating the Vo part from the remaining vocal parts, referred to as Vo separation, enables downstream applications such as lyric recognition and minus-one accompaniment generation for vocal ensemble music. Despite these potential applications, acoustic cues for this task are limited because the target and interfering sources are all singing voices with similar acoustic characteristics and often overlap in time, making Vo separation challenging. In this paper, we propose a Vo separation model that uses phoneme alignment of the Vo part as auxiliary information. The proposed model is based on band-split RoPE Transformer (BS-RoFormer), a state-of-the-art music source separation model, and introduces frame-level phoneme labels into its intermediate representations using feature-wise linear modulation (FiLM). Experimental results show that phoneme-alignment conditioning improves Vo separation performance over an audio-only baseline and yields larger average gains than conditioning only on Vo singing/silence activity. Further analysis suggests that the advantage of phoneme-label information is larger when fewer remaining vocal parts share the same phoneme as Vo.
Problem

Research questions and friction points this paper is trying to address.

Lead Vocal Separation
Vocal Ensemble Mixtures
Phoneme Alignment
Music Source Separation
Innovation

Methods, ideas, or system contributions that make the work stand out.

phoneme alignment
frame-level phoneme labels
feature-wise linear modulation (FiLM)
band-split RoPE Transformer (BS-RoFormer)
🔎 Similar Papers
No similar papers found.