Voice of Reason: Reinforcement Learning for Spoken Math

📅 2026-09-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该研究通过强化学习提高语音模型在数学推理上的准确性,使用GLM-4-Voice模型并结合监督微调和现有流式推理技术,达到74.8%的自由形式准确率。
📝 Abstract
Speech language models enable richer spoken interactions between humans and machines than cascaded systems, allowing access to paralinguistic information and lower latency. However, their accuracy on mathematical reasoning benchmarks has lagged behind those of text models. Reinforcement learning (RL) with verifiable rewards has been instrumental in extending text models' capabilities for solving complex problems and limiting hallucinations. In this work, we explore applying RL to the GLM-4-Voice speech model (Zeng et al., 2024) to bridge the gap between textual and spoken mathematical problem solving. We first adapt the model to the domain using supervised fine-tuning on synthesized spoken question-answering data. We then show that, even without extra reasoning tokens, RL improves the accuracy on GSM8K beyond levels previously achieved for speech models only with supplementary reasoning traces. When combined with existing streaming reasoning techniques, we show further gains to 74.8% free-form accuracy. This establishes a new state-of-the-art for mathematical spoken abilities with speech-native models.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning
Spoken Math
Accuracy
Mathematical Reasoning
Speech Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement Learning
Speech Models
Mathematical Reasoning
GSM8K
Streaming Reasoning