VIOLET: High-Fidelity Violin Synthesis with Techniques and Dynamics

📅 2026-08-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing neural audio synthesis approaches predominantly focus on piano and struggle to faithfully model the expressive techniques and continuous dynamics of sustained-tone instruments such as the violin. To address this limitation, this work proposes VIOLET—the first end-to-end controllable neural synthesis system for violin—built upon a latent diffusion framework that integrates a Diffusion Transformer with rectified flow algorithms. VIOLET generates high-fidelity audio conditioned on MIDI notes, expressive technique labels, and continuous dynamic contours. The study also introduces CSV-TD, a newly curated dataset comprising 39 hours of meticulously annotated violin performances. Experimental results demonstrate that VIOLET significantly outperforms existing methods in terms of articulation accuracy, pitch-temporal alignment, and dynamic controllability, achieving performance comparable to state-of-the-art commercial virtual instruments.
📝 Abstract
Neural synthesis for musical instruments has the potential to revolutionize current practices that use concatenative synthesis and a sample library. However, most research focused on piano synthesis and expressive performance generation; little work has been done on continuously articulated instruments like the violin, let alone rendering them with playing techniques and dynamics. We present VIOLET, a latent-diffusion framework for controllable violin synthesis, which uses a Diffusion Transformer (DiT) with rectified flow to synthesize high-fidelity audio from MIDI notes, playing techniques, and continuous dynamics. To train VIOLET, in addition to using a few existing datasets, we curate a new dataset named CSV-TD, which contains 39 h of 48 kHz synthetic audio and time-aligned annotations of MIDI notes, note-level techniques, and continuous dynamics curves. Objective and subjective evaluations show that VIOLET synthesizes violin performances with high technique adherence, accurate pitch and timing alignment, and good dynamics control. It outperforms the current state-of-the-art neural violin synthesis system and approaches a top commercial virtual instrument in terms of technique clarity, naturalness, and dynamics following.
Problem

Research questions and friction points this paper is trying to address.

violin synthesis
playing techniques
dynamics
neural audio synthesis
expressive performance
Innovation

Methods, ideas, or system contributions that make the work stand out.

latent diffusion
violin synthesis
playing techniques
continuous dynamics
Diffusion Transformer
💼 Related Jobs
No related jobs found.
Baotong Tian
Baotong Tian
Ph.D. student at University of Rochester
Machine LearningExplainable AIMusic Information RetrievalMusic Generation
C
Cynthia Lu
University of Rochester, USA
V
Vincent K. M. Cheung
Sony Computer Science Laboratories, Tokyo, Japan
T
Ting-Kang Wang
National Taiwan University, Taiwan
J
Jonathan Churchill
Embertone, USA
Zhiyao Duan
Zhiyao Duan
Professor of Electrical and Computer Engineering, University of Rochester
Computer AuditionMusic Information RetrievalSpeech ProcessingAudiovisual Learning