SPHERE: Automatic Music Upmixing via Audio Language Model Post-Training with Spatial Heuristic Rewards

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文研究自动音乐上混任务,通过音频语言模型后训练和空间启发式奖励方法,预测多音轨录音的空间混合参数。
📝 Abstract
In this paper, we study the task of automatic music upmixing, wherein a system predicts spatial mixing parameters from a multi-stem recording. Different from existing methods that rely on task-specific music encoders, we approach this task via audio language model (ALM) post-training, leveraging rich representations from existing ALMs, which encode both music semantics and mixing knowledge. Specifically, we propose a post-training recipe that first employs rejection sampling SFT, followed by reinforcement learning (RL) with verifiable rewards (RLVR) via GRPO. We propose Sphere (Spatial Heuristic Rewards), a deterministic reward suite inspired by music mixing conventions, to guide our post-training. It consists of 6 perceptually-motivated sub-rewards and encourages the output mix to be centered, balanced and spacious. More broadly, our results suggest that expert domain knowledge can be encoded as verifiable rewards and distilled into language models, without task-specific architectures.
Problem

Research questions and friction points this paper is trying to address.

automatic music upmixing
multi-stem recording
spatial mixing parameters
Innovation

Methods, ideas, or system contributions that make the work stand out.

Audio Language Model
Post-Training
Spatial Heuristic Rewards
Reinforcement Learning with Verifiable Rewards (RLVR)
Automatic Music Upmixing