Voice of Reason: Reinforcement Learning for Spoken Math
该研究通过强化学习提高语音模型在数学推理上的准确性,使用GLM-4-Voice模型并结合监督微调和现有流式推理技术,达到74.8%的自由形式准确率。
该研究通过强化学习提高语音模型在数学推理上的准确性,使用GLM-4-Voice模型并结合监督微调和现有流式推理技术,达到74.8%的自由形式准确率。
本文针对多语言大模型的词汇压缩不均和资源浪费问题,提出模块化分词器学习方法及预训练策略,提高效率与跨语言公平性。
This work addresses the unnatural interaction behaviors—such as excessively long silences and improper turn-taking—that commonly arise in full-duplex spoken dialogue systems due to their reliance on token-level likelihood optimization. To overcome these limitations, the authors propose a reinforcement learning–based post-training alignment method that systematically encompasses four key interactive dimensions: pause handling, turn-taking, backchannel feedback, and user interruption. The approach introduces an audio segment–level reward function and incorporates a large language model to constrain semantic quality. Experimental results on the Moshi and PersonaPlex benchmarks demonstrate that the method significantly enhances both offline and real-time multi-turn dialogue fluency and naturalness while preserving response accuracy.
This work investigates how the temporal ordering of training data affects large language models’ ability to acquire time-sensitive knowledge. Training a 6B-parameter model on chronologically ordered Common Crawl snapshots, the study systematically examines whether sequential exposure to temporally structured corpora mitigates knowledge freezing—a common issue in models trained on shuffled data that impairs their capacity to accurately associate facts with their occurrence times. The authors introduce a novel evaluation benchmark comprising over 7,000 temporally grounded questions and provide the first empirical evidence that temporally ordered pretraining significantly enhances model performance on fact freshness and temporal accuracy, without compromising general language understanding capabilities compared to conventional shuffled-data training.
This work addresses the high computational cost in large language model (LLM) inference caused by contextual redundancy. We propose ARC-Encoder, a general-purpose, architecture- and parameter-agnostic context compression method that requires no modification to the target LLM. ARC-Encoder learns continuous, compact textual representations to replace original token embeddings, directly injecting them into the decoder’s input layer. Its core contribution is a unified encoder architecture—designed for broad compatibility across diverse LLM families—combined with continuous representation learning and a systematic training strategy, enabling 4×–8× context compression without fine-tuning the target model. Experiments demonstrate significant reductions in inference latency and GPU memory consumption across both instruction-tuned and base LLMs, with seamless plug-and-play deployment. ARC-Encoder achieves state-of-the-art performance in efficiency and practicality.
该研究通过强化学习提高语音模型在数学推理上的准确性,使用GLM-4-Voice模型并结合监督微调和现有流式推理技术,达到74.8%的自由形式准确率。
本文针对多语言大模型的词汇压缩不均和资源浪费问题,提出模块化分词器学习方法及预训练策略,提高效率与跨语言公平性。
This work addresses the unnatural interaction behaviors—such as excessively long silences and improper turn-taking—that commonly arise in full-duplex spoken dialogue systems due to their reliance on token-level likelihood optimization. To overcome these limitations, the authors propose a reinforcement learning–based post-training alignment method that systematically encompasses four key interactive dimensions: pause handling, turn-taking, backchannel feedback, and user interruption. The approach introduces an audio segment–level reward function and incorporates a large language model to constrain semantic quality. Experimental results on the Moshi and PersonaPlex benchmarks demonstrate that the method significantly enhances both offline and real-time multi-turn dialogue fluency and naturalness while preserving response accuracy.
This work investigates how the temporal ordering of training data affects large language models’ ability to acquire time-sensitive knowledge. Training a 6B-parameter model on chronologically ordered Common Crawl snapshots, the study systematically examines whether sequential exposure to temporally structured corpora mitigates knowledge freezing—a common issue in models trained on shuffled data that impairs their capacity to accurately associate facts with their occurrence times. The authors introduce a novel evaluation benchmark comprising over 7,000 temporally grounded questions and provide the first empirical evidence that temporally ordered pretraining significantly enhances model performance on fact freshness and temporal accuracy, without compromising general language understanding capabilities compared to conventional shuffled-data training.
This work addresses the high computational cost in large language model (LLM) inference caused by contextual redundancy. We propose ARC-Encoder, a general-purpose, architecture- and parameter-agnostic context compression method that requires no modification to the target LLM. ARC-Encoder learns continuous, compact textual representations to replace original token embeddings, directly injecting them into the decoder’s input layer. Its core contribution is a unified encoder architecture—designed for broad compatibility across diverse LLM families—combined with continuous representation learning and a systematic training strategy, enabling 4×–8× context compression without fine-tuning the target model. Experiments demonstrate significant reductions in inference latency and GPU memory consumption across both instruction-tuned and base LLMs, with seamless plug-and-play deployment. ARC-Encoder achieves state-of-the-art performance in efficiency and practicality.