TART: A Modular Tool for Technique-Aware Audio-to-Tablature Guitar Transcription

📅 2026-09-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决吉他自动音乐转录中表达技巧捕捉、弦品组合错误及噪声环境适应性问题,提出TART工具,通过四阶段流程实现音频到谱表的转换。
📝 Abstract
Automatic Music Transcription (AMT) for guitar remains limited by three challenges: existing systems often fail to capture expressive techniques such as slides, bends, and percussive hits; they often assign notes to incorrect string-fret combinations; and they are typically trained on clean recordings, limiting their generalization to noisy real-world audio. To address these challenges, we propose TART, a modular four-stage audio-to-tablature pipeline consisting of (1) an audio-to-MIDI transcription model, (2) an expressive technique classifier, (3) an audio-conditioned T5 encoder-decoder for string-fret assignment, and (4) an automated tablature generator. We evaluate TART in a zero-shot setting on GuitarSet, EGDB, and two augmented benchmarks, Noisy GuitarSet and Noisy EGDB. Averaged across these four benchmarks, TART achieves 81.35% audio-to-MIDI F50 (+6.67 points over the best prior baseline), 71.8% string-fret Tab F1 (+8.5 points over the best prior baseline), and 54.08% end-to-end Tab F1. To our knowledge, TART is the first framework to generate guitar tablature with both fingering and expressive technique annotations directly from guitar audio.
Problem

Research questions and friction points this paper is trying to address.

Automatic Music Transcription
expressive techniques
string-fret assignment
noisy real-world audio
Innovation

Methods, ideas, or system contributions that make the work stand out.

Technique-Aware
Audio-to-Tablature
Modular Pipeline
Expressive Techniques
String-Fret Assignment
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Akshaj Gupta
University of California, Berkeley
H
Hwi Joo Park
University of California, Berkeley
A
Andrea Guzman
University of California, Berkeley
S
Shamak Gowda
University of California, Berkeley
S
Samhita Konduri
University of California, Berkeley
Jiachen Lian
Jiachen Lian
UC Berkeley
precision healthcarespeech processingmachine learning
Robin Netzorg
Robin Netzorg
PhD Student, UC Berkeley
Speech ModificationInterpretability
G
Gopala Anumanchipalli
University of California, Berkeley