Preference Optimization with LALM Feedback for Continuous Autoregressive Non-Verbal Vocalization Generation

📅 2026-09-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出了一种利用大音频-语言模型反馈的偏好优化框架,用于可控的非言语发声生成,并通过拒绝采样微调和锚定流-DPO两阶段优化策略来提高生成质量。
📝 Abstract
We propose a preference optimization framework with Large Audio-Language Model (LALM) feedback for controllable non-verbal vocalization (NVV) generation in continuous autoregressive speech models. To construct preference data without human preference annotation, we build a bilingual prompt corpus by combining NVV-injected real transcripts with LLM-generated semantically aligned prompts, perform stochastic model rollouts, and use a LALM to rank candidate utterances and form same-prompt chosen--rejected pairs. We then adopt a two-stage optimization strategy: Rejection Sampling Fine-Tuning (RSFT) first adapts the model to LALM-selected high-scoring samples, followed by Anchored Flow-DPO, which formulates pairwise preference optimization using utterance-level flow-matching loss and retains the chosen-sample flow-matching objective as an SFT anchor. This design enables DPO-style preference learning without explicit sequence likelihoods while preserving direct supervision on preferred realizations. On the official 1,600-utterance NVVSpeech Challenge Track~2 test set, our method achieves a Final Track2Score of \textbf{75.80} (79.39 ZH / 72.21 EN), outperforming the VoxCPM2 baseline by \textbf{+1.84}. The improvements are mainly driven by higher NVV Accuracy and NVV Perceptual Effect, while Overall Quality remains stable.
Problem

Research questions and friction points this paper is trying to address.

Non-Verbal Vocalization
Continuous Autoregressive Speech Models
Preference Optimization
LALM Feedback
Controllable Generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Preference Optimization
Large Audio-Language Model (LALM) Feedback
Non-Verbal Vocalization (NVV) Generation
Rejection Sampling Fine-Tuning (RSFT)
Anchored Flow-DPO
🔎 Similar Papers
No similar papers found.
J
Jingbin Hu
Audio, Speech and Language Processing Group (ASLP@NPU), School of Computer Science, Northwestern Polytechnical University, Xi’an, China
Q
Qirui Zhan
Audio, Speech and Language Processing Group (ASLP@NPU), School of Computer Science, Northwestern Polytechnical University, Xi’an, China
Y
Yuang Cao
Audio, Speech and Language Processing Group (ASLP@NPU), School of Computer Science, Northwestern Polytechnical University, Xi’an, China
Z
Ziyu Zhang
Audio, Speech and Language Processing Group (ASLP@NPU), School of Computer Science, Northwestern Polytechnical University, Xi’an, China
Y
Yunxiang Chen
Shenzhen Pimei Technology Co., Ltd., Guangdong, China
H
Houdun Liu
Shenzhen Pimei Technology Co., Ltd., Guangdong, China
S
Shuo Feng
Shenzhen Pimei Technology Co., Ltd., Guangdong, China
B
Bengu Wu
Yutu Zhineng, Beijing, China
Lei Xie
Lei Xie
Northwestern Polytechnical University
speech processingspeech recognitionspeech synthesismultimediaartificial intelligence
Liumeng Xue
Liumeng Xue
Hong Kong University of Science and Technology
Audio Speech and Language ProcessingSpeech Generation