Discriminative Flow Matching: Beyond Time-Conditioning in Generative Restoration via Flow-State Representations
本文提出Discriminative Flow Matching方法,通过使用判别表示而非时间坐标来解决生成恢复中样本依赖性的问题。
本文提出Discriminative Flow Matching方法,通过使用判别表示而非时间坐标来解决生成恢复中样本依赖性的问题。
This study addresses the challenge posed by imprecise note onset annotations in weakly aligned score–audio data, which significantly limits the performance of automatic music transcription. It presents the first systematic analysis of the impact of onset “snapping” during cross-instrument transcription training and introduces a global context–aware optimization method. The approach formulates snapping as a pitch-wise assignment problem, constructing a bipartite graph from neural network posteriorgrams and dynamic time warping alignments, and replaces conventional greedy strategies with optimal matching within context-sensitive temporal windows. Evaluated on piano, chamber, and orchestral datasets, the method substantially improves both onset alignment accuracy and overall transcription performance, with particularly pronounced gains under wide snapping windows or coarse initial alignments.
Traditional pitch tracking methods exhibit degraded performance on non-music and non-speech audio—such as bioacoustic and environmental recordings—due to wide bandwidths and rapidly varying fundamental frequencies. To address this, we propose a novel paradigm for pitch contour analysis that bypasses explicit pitch tracking altogether; instead, we model pitch contours as structured visual patterns in time–frequency spectrograms. Leveraging transfer learning, we adapt a natural-image pre-trained object detection architecture (YOLOv5), fine-tuned on synthetically generated pitch contour data, to directly localize and regress pitch trajectories within spectrograms. Evaluated across eight downstream tasks spanning music, speech, bioacoustics, and environmental audio, our method consistently outperforms state-of-the-art pitch trackers—including CREPE and PYIN—demonstrating superior cross-domain generalization and robustness to spectral variability and rapid pitch modulation.
本文提出Discriminative Flow Matching方法,通过使用判别表示而非时间坐标来解决生成恢复中样本依赖性的问题。
This study addresses the challenge posed by imprecise note onset annotations in weakly aligned score–audio data, which significantly limits the performance of automatic music transcription. It presents the first systematic analysis of the impact of onset “snapping” during cross-instrument transcription training and introduces a global context–aware optimization method. The approach formulates snapping as a pitch-wise assignment problem, constructing a bipartite graph from neural network posteriorgrams and dynamic time warping alignments, and replaces conventional greedy strategies with optimal matching within context-sensitive temporal windows. Evaluated on piano, chamber, and orchestral datasets, the method substantially improves both onset alignment accuracy and overall transcription performance, with particularly pronounced gains under wide snapping windows or coarse initial alignments.
Traditional pitch tracking methods exhibit degraded performance on non-music and non-speech audio—such as bioacoustic and environmental recordings—due to wide bandwidths and rapidly varying fundamental frequencies. To address this, we propose a novel paradigm for pitch contour analysis that bypasses explicit pitch tracking altogether; instead, we model pitch contours as structured visual patterns in time–frequency spectrograms. Leveraging transfer learning, we adapt a natural-image pre-trained object detection architecture (YOLOv5), fine-tuned on synthetically generated pitch contour data, to directly localize and regress pitch trajectories within spectrograms. Evaluated across eight downstream tasks spanning music, speech, bioacoustics, and environmental audio, our method consistently outperforms state-of-the-art pitch trackers—including CREPE and PYIN—demonstrating superior cross-domain generalization and robustness to spectral variability and rapid pitch modulation.