H2H Music Improv: A Communication Model and Audio-Visual Dataset for Music Improvisation

📅 2026-08-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the critical bottleneck in real-time AI improvisation systems arising from their lack of human communicative awareness and formalized communication models. To bridge this gap, we propose the first machine-readable framework for free improvisation communication. Leveraging expert co-creation and multimodal acquisition techniques, we constructed a six-hour audiovisual dataset featuring isolated audio tracks and bidirectional intention annotations from expert duos. This work fills a significant void in the field by providing both a rigorous theoretical foundation for analyzing musician interaction mechanisms and high-quality data resources. Ultimately, these contributions are instrumental for developing AI improvisational partners endowed with native communicative capabilities, thereby advancing human-AI musical collaboration.
📝 Abstract
Current real-time AI improvisation systems lack the communication awareness human musicians rely on: rather than treating communication as a foundational algorithm design concern, most systems layer interaction strategies post-hoc onto generative algorithms through explicit controls and predefined modes. This gap persists in part because no formalized, machine-readable communication model with musicians exists. To address this, we study how expert musicians communicate in free (non-idiomatic) improvisation, unconstrained by prior discussion or agreement. Through a collaborative co-design process with expert improvisers, we derive a communication model that (1) captures how free improvisers negotiate musical ideas and enter stable musical spaces, and (2) is formalized as a machine-readable annotation scheme. We further present the H2H (Human-to-Human) Music Improvisation dataset: six hours of audio-visual expert duo improvisations with clean per-player stems and per-player annotations of both their own intentions and their perception of their partner's intentions. To our knowledge, this is the first such dataset for free improvisation. Together, the communication model and the dataset offer a new lens and resource for studying musician communication and may in future inform the design of AI musical partners that communicate by design.
Problem

Research questions and friction points this paper is trying to address.

AI music improvisation
musician communication
free improvisation
communication model
audio-visual dataset
Innovation

Methods, ideas, or system contributions that make the work stand out.

Communication Model
Free Improvisation
Audio-Visual Dataset
Intent Annotation
Human-to-Human Interaction
🔎 Similar Papers
No similar papers found.
A
Aleksandra Teng Ma
Human-AI Resonance Lab, Massachusetts Institute of Technology, USA
A
Anthony Cammarota
Creative Music Technology Lab, Georgia Institute of Technology, USA
Jiayi Wang
Jiayi Wang
Music Informatics Group, Georgia Institute of Technology, USA
A
Alexandria Smith
Creative Music Technology Lab, Georgia Institute of Technology, USA
Cheng-Zhi Anna Huang
Cheng-Zhi Anna Huang
MIT Music and Theater Arts (MTA) and Electrical Engineering and Computer Science (EECS)
Music GenerationDeep LearningHuman-Computer InteractionCo-Creativity
J
Jeffrey Albert
Creative Music Technology Lab, Georgia Institute of Technology, USA
Alexander Lerch
Alexander Lerch
Music Informatics Group, Georgia Institute of Technology
audio content analysismusic information retrievalsemantic audioaudio signal processingmusic generation