SwanWeave:One-Stage Multi-Task Instruction-Guided 3D Spatial Audio Editing

📅 2026-09-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出SwanWeave,一种基于指令的一阶段多任务3D空间音频编辑框架,使用SE-MoE和SPO方法解决复杂3D空间音频编辑问题。
📝 Abstract
Spatial audio editing modifies an existing soundfield according to a user's instruction while preserving the rest of the scene. Unlike conventional audio editing, it must reason jointly about audio events, spatial information, dynamic changes, and environmental information in first-order Ambisonic (FOA) waveforms. Existing language-guided editors mainly target conventional audio or rely on sequential operations, and therefore do not directly support one-stage editing for complex 3D spatial instructions. We present SwanWeave, the first one-stage multi-task framework for instruction-guided 3D FOA spatial audio editing. We build paired FOA supervision from open-source speech and sound-effect corpora using controllable room simulation, covering more than ten single-operation and compound tasks across the four editing axes. To handle this heterogeneous edit space, SwanWeave uses Spatial Edit Mixture-of-Experts (SE-MoE) with dual-level routing, selecting task-aware expert combinations for compound instructions and frame-level routed/null experts for local edit decisions. We further introduce Spatial Preference Optimization (SPO), a Direct Preference Optimization (DPO)-based alignment objective with edit-specific negative targets, and adopt staged training to improve natural-language grounding. Experiments show that SwanWeave achieves better editing quality than existing general audio editors and spatial audio baselines across all tasks. Spatial audio editing demos can be found at https://swanaigc.github.io/#swanweave, code can be found at: https://github.com/MM-Speech/SwanWeave.
Problem

Research questions and friction points this paper is trying to address.

spatial audio editing
instruction-guided
first-order Ambisonic (FOA)
one-stage multi-task
Innovation

Methods, ideas, or system contributions that make the work stand out.

one-stage multi-task
FOA spatial audio editing
Spatial Edit Mixture-of-Experts (SE-MoE)
Spatial Preference Optimization (SPO)