MazzikaAI: A knowledge-based performance-to-prompt compiler for real-time Arabic maqam accompaniment with a streaming text-to-music model

📅 2026-08-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of existing generative music models, which predominantly rely on Western twelve-tone equal temperament and thus struggle to support the microtonal intervals, ornamental nuances, and real-time interactive accompaniment required by Arabic maqam music. The authors propose a knowledge-driven real-time accompaniment framework that dynamically compiles performer inputs—such as MIDI, gestures, and harmonic context—into natural language prompts via an expert rule base, directly guiding an unmodified, streaming text-to-music foundation model (Google Lyria RealTime) to generate culturally appropriate accompaniments. By integrating deterministic musicological rules with a general-purpose generative model without fine-tuning, this approach achieves sub-second end-to-end latency, significantly increases the usage of quarter tones, and enhances maqam fidelity, thereby demonstrating the feasibility of culturally inclusive audio generation.
📝 Abstract
Arabic maqam music microtonal, modal, and built on ornamented call and response is among the traditions most underserved by generative music models, whose training frameworks remain predominantly Western and equaltempered. Real time accompaniment sharpens this gap: an AI partner must listen, adapt dynamically, and respect idiomatic microtonal structures. Streaming text to music models provide strong generative capabilities but lack precise control interfaces. We present MazzikaAI, a knowledge based system that uses natural language as the actuator of a realtime control loop. By compiling live MIDI, gesture, and inferred harmony into continuously updated text prompts, MazzikaAI steers an unmodified streaming generator, Google Lyria RealTime, without requiring model finetuning. The system embeds expert knowledge of six core maqamat, characteristic ornaments, and ensemble dynamics, maintaining realtime responsiveness with subsecond keytoaudibleupdate latency. Empirical evaluations demonstrate that dynamic prompt compilation reliably grounds generation in microtonal scales, significantly increasing offgrid quartertone content over baseline generation. Beyond its core implementation, MazzikaAI illustrates how deterministic knowledgebased rules can effectively bridge expert, nonWestern musical traditions and unfinetuned foundation models. This architecture establishes a scalable paradigm for realtime humanAI cocreation, offering a generalizable blueprint for interactive accompaniment, adaptive music education, and culturally inclusive generative audio across diverse global idioms.
Problem

Research questions and friction points this paper is trying to address.

Arabic maqam
microtonal music
real-time accompaniment
generative music models
non-Western musical traditions
Innovation

Methods, ideas, or system contributions that make the work stand out.

knowledge-based compilation
real-time accompaniment
microtonal control
prompt engineering
non-Western music generation