Direct Preference Density Alignment for Conversational Audio Equalization

📅 2026-09-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
提出直接偏好密度对齐方法,利用大规模用户数据构建非参数偏好密度图,结合在线和离线优化优势,解决对话音频均衡问题。
📝 Abstract
Large Language Model alignment typically relies on learned proxy reward models, which significantly increase the memory footprint during training and are notoriously prone to instability and reward hacking. While offline methods like Direct Preference Optimization (DPO) bypass the reward model, they lose the ability to perform online exploration. If no optimization constraints are applied, this can lead to format collapse in bounded, continuous spaces. To resolve this, we propose Direct Preference Density Alignment: An alternative framework that removes the need for a learned proxy reward model while strictly preserving the benefits of online reinforcement learning. We leverage large-scale user data (approximately 90,000 samples) to construct non-parametric preference density maps, establishing an empirical reward surface. In addition to removing the reward model, Direct Preference Density Alignment enables the combination of the online structural grounding of Group Relative Policy Optimization (GRPO) with the targeted offline refinement of DPO. We show that this GRPO+DPO combination achieves the highest performance, and in a blind audio equalization listening test, enables a 1.5B-parameter model to achieve perceptual parity with a carefully prompt-engineered GPT-4o mini baseline, using only a fraction of the inference compute.
Problem

Research questions and friction points this paper is trying to address.

Large Language Model
reward model
Direct Preference Optimization
online exploration
format collapse
Innovation

Methods, ideas, or system contributions that make the work stand out.

Direct Preference Density Alignment
non-parametric preference density maps
Group Relative Policy Optimization (GRPO)
Direct Preference Optimization (DPO)
🔎 Similar Papers
No similar papers found.