CANTABILE: Learning Expressive Dynamics for Robotic Piano Performance

📅 2026-09-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究提出CANTABILE框架,通过结合速度目标和奖励机制改进机器人钢琴演奏的动态表现力,显著提高音符准确度和音乐表现力。
📝 Abstract
Robotic piano playing has emerged as a standard benchmark for dexterous bimanual manipulation, yet progress on it has been measured almost entirely by note accuracy -- which keys are pressed (pitch) and when (onset) -- leaving the musical dynamics essential for expressive performance neither rewarded nor evaluated. We propose CANTABILE, a dynamics-aware framework for robotic piano performance that (i) closes the score-to-contact loop by conditioning the policy on upcoming velocity goals and mapping each key's angular velocity at onset back to MIDI velocity, (ii) couples a velocity-fidelity reward with an onset-coverage reward, so that dynamics cannot be improved by omitting difficult notes, and (iii) refines a frozen dynamics-aware base policy with an alpha-scaled, finger-only residual that localizes strike-intensity adaptation away from nominal note execution. On EXPRESSIVE-51, a dynamics-rich 51-song subset of RoboPianist, CANTABILE raises Velocity F1 -- jointly measuring pitch, onset, and intensity within a +/-8 MIDI-velocity tolerance -- from 0.06 to 0.34 over the RoboPianist baseline, improves all 51 songs, more than halves matched-note velocity error, and reduces log-mel distance to reference audio by 8%. Intensity-randomized training further enables runtime control of performance intensity without retraining.
Problem

Research questions and friction points this paper is trying to address.

robotic piano
dynamics
expressive performance
Innovation

Methods, ideas, or system contributions that make the work stand out.

dynamics-aware framework
velocity-fidelity reward
onset-coverage reward
residual adaptation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
W
Woosik Kim
KAIST, Daejeon, Republic of Korea
W
Wonhyeok Choi
DGIST, Daegu, Republic of Korea
Sunghoon Im
Sunghoon Im
EECS, DGIST
Computer VisionDeep LearningRobot Vision