TUTTI: Toward generalizable audio-to-score transcription via fully synthesized data
为解决高质量音频到乐谱转录数据稀缺问题,研究通过完全合成的数据和Transformer模型预训练方法,提高跨乐器转录性能。
为解决高质量音频到乐谱转录数据稀缺问题,研究通过完全合成的数据和Transformer模型预训练方法,提高跨乐器转录性能。
Traditional approaches struggle to jointly model natural speech and controllable singing within a single framework due to the fundamentally different constraints imposed by prosody and melody. This work proposes UniVoice, a unified generative architecture based on conditional flow matching that decomposes conditioning signals into content, melody, and timbre, and employs a shared DiT backbone for multimodal fusion. A key innovation is the introduction of learnable null melody tokens to replace explicit melody inputs during speech synthesis, thereby preserving explicit melodic control for singing while avoiding redundant constraints on speech; theoretical analysis shows this design approximates marginalization over melody. Trained on 30k hours of speech and 35k hours of singing data, UniVoice achieves a speech PER of 5.26%, comparable to specialized TTS systems, and a singing PER of 16.22%, substantially outperforming the unified baseline Vevo1.5 (24.72%).
Aesthetic evaluation of generated music remains challenging due to the complexity of perceptual dimensions. To address this, we propose a multi-scale hierarchical evaluation framework: (1) a cross-paragraph attention mechanism jointly models local musical details and global structural coherence, integrated with a multi-scale convolutional network; (2) a semantics-preserving C-Mixup audio augmentation strategy enhances data diversity and model robustness; and (3) a regression-ranking joint optimization objective enables consistent learning across segment-level score prediction and full-track ranking. Evaluated on the ICASSP 2026 SongEval benchmark, our method significantly outperforms existing baselines—achieving a 12.3% improvement in Pearson correlation coefficient and a 9.7% gain in Top-10 high-quality song identification accuracy. To our knowledge, this is the first approach to effectively balance multidimensional aesthetic consistency with end-to-end trainability.
为解决高质量音频到乐谱转录数据稀缺问题,研究通过完全合成的数据和Transformer模型预训练方法,提高跨乐器转录性能。
Traditional approaches struggle to jointly model natural speech and controllable singing within a single framework due to the fundamentally different constraints imposed by prosody and melody. This work proposes UniVoice, a unified generative architecture based on conditional flow matching that decomposes conditioning signals into content, melody, and timbre, and employs a shared DiT backbone for multimodal fusion. A key innovation is the introduction of learnable null melody tokens to replace explicit melody inputs during speech synthesis, thereby preserving explicit melodic control for singing while avoiding redundant constraints on speech; theoretical analysis shows this design approximates marginalization over melody. Trained on 30k hours of speech and 35k hours of singing data, UniVoice achieves a speech PER of 5.26%, comparable to specialized TTS systems, and a singing PER of 16.22%, substantially outperforming the unified baseline Vevo1.5 (24.72%).
Aesthetic evaluation of generated music remains challenging due to the complexity of perceptual dimensions. To address this, we propose a multi-scale hierarchical evaluation framework: (1) a cross-paragraph attention mechanism jointly models local musical details and global structural coherence, integrated with a multi-scale convolutional network; (2) a semantics-preserving C-Mixup audio augmentation strategy enhances data diversity and model robustness; and (3) a regression-ranking joint optimization objective enables consistent learning across segment-level score prediction and full-track ranking. Evaluated on the ICASSP 2026 SongEval benchmark, our method significantly outperforms existing baselines—achieving a 12.3% improvement in Pearson correlation coefficient and a 9.7% gain in Top-10 high-quality song identification accuracy. To our knowledge, this is the first approach to effectively balance multidimensional aesthetic consistency with end-to-end trainability.