Agogic: Performance-Timed Music Tokens for LLM-Native Text-to-Symbolic-Music Generation

📅 2026-08-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of systematic evaluation of music tokenization’s impact in text-to-symbolic-music generation. By fixing the pretrained model, training data, and decoding strategy while varying only the symbolic representation, the authors systematically assess seven tokenization schemes in terms of distributional fidelity. Their findings reveal that the choice of musical representation—not model scale—is the primary determinant of generation quality. They propose PMT (Performance-level Music Tokenization), a high-fidelity representation that enables even a 0.8B-parameter model to achieve an FMD score of 159, substantially outperforming 27B-parameter models using beat-grid representations (FMD: 272–286). Additionally, lightweight decoding constraints double instrument F1 and tonality accuracy while preserving distributional consistency.
📝 Abstract
Text-to-music language models begin with a choice usually made by default: how to tokenize music. Normally entangled with backbone, data, and recipe, its effect has never been measured in isolation. We fix pretrained Qwen3.5 (0.8B-27B), data, budget, and decoding, and swap only the representation across seven tokenizations, anchoring texture metrics to each representation's model-free ceiling. The ordering is clean and surprising: representation, not model size, is the binding variable for distributional fidelity. Scaling the backbone 34x barely moves Frechet Music Distance (FMD), whereas switching representation halves it. PMT, a performance-resolution stream we release (10 ms timing, per-note velocity, multi-track texture; 609 symbols), reaches FMD 159 at 0.8B against 272-286 for beat grids (1.7-1.8x lower, up to 2.8x elsewhere; non-overlapping bootstrap CIs), so a 0.8B performance-resolution model beats a 27B beat grid. It reappears on a 26M from-scratch backbone and a second performance-resolution tokenizer: a property of the class, not one lucky vocabulary. Nor is it a finer-lattice artifact: snapping PMT's onsets to the beat grids' resolution still leaves it 67-129 FMD ahead of both (n=500). The effect is distributional; whether it is audible is a separate question, left open by our probe, with a human study pre-registered. Native caption adherence is weak but separable: a lightweight decode-time constraint doubles instrument-F1 (.28 to .60) and Correct-Key (.16 to .35) at no distributional cost. We release the harness, 25+ checkpoints, two corpora (86.6k aligned across caption/MIDI/ABC/audio; 6.25M captioned, the largest for music), and an imprinting diagnostic: published text-to-MIDI systems reproduce their training distribution near-invariant to the caption (72% vs. 71% chord-time on disjoint domains). The field's next representation claim can now be measured, not asserted.
Problem

Research questions and friction points this paper is trying to address.

music tokenization
text-to-music generation
representation
distributional fidelity
symbolic music
Innovation

Methods, ideas, or system contributions that make the work stand out.

music tokenization
performance-timed tokens
text-to-music generation
representation evaluation
Frechet Music Distance
🔎 Similar Papers
No similar papers found.
J
Junhao Chen
Tsinghua University
M
Mingjin Chen
The Hong Kong Polytechnic University
J
Jingjia Mao
Tsinghua University
L
Lin Chen
Beijing Technology and Business University
Saining Zhang
Saining Zhang
College of Computing and Data Science, Nanyang Technological University
Computer Vision
Minglin Chen
Minglin Chen
Sun Yat-sen University, Ph.D. Student
3D VisionComputer VisionComputer GraphicDeep Learning
R
Ruocheng Wu
The University of Hong Kong
L
Liaoyuan Fan
The University of Hong Kong
W
Wenyi Li
University of Chinese Academy of Sciences
Mingju Gao
Mingju Gao
Unknown affiliation
Computer VisionRobotics
H
Henghaofan Zhang
University of Electronic Science and Technology of China
Z
Zhihao Li
SparcAI Inc.
Hao Zhao
Hao Zhao
Tsinghua University
Computer Vision
Y
Yufei Wang
SparcAI Inc.
Ruqi Huang
Ruqi Huang
Tsinghua Shenzhen International Graduate School
3D Computer VisionShape AnalysisGeometry Processing