ECHO: A Matched-Contrast Benchmark for Context-Sensitive Turn-Taking in Full-Duplex Dialogue
研究解决了全双工对话系统中区分打断和反馈的问题,通过创建ECHO基准测试,使用匹配对比方法评估上下文敏感的决策。
研究解决了全双工对话系统中区分打断和反馈的问题,通过创建ECHO基准测试,使用匹配对比方法评估上下文敏感的决策。
为解决全双工对话模型数据不足的问题,通过构建包含情景、行为、情感和声音事件的DuplexDrama合成对话数据集来提供支持。
为解决多轮代理任务中稀疏终端奖励导致的中间动作信用分配不精细问题,提出基于势能引导的策略优化方法PGPO,通过估计状态势能并传播跨轨迹信用。
Flow Matching in speech synthesis suffers from high inference latency and timbre leakage. This work proposes a unified guidance framework that, for the first time, jointly integrates data-level and model-level guidance. By leveraging heterogeneous data augmentation to disentangle linguistic content from acoustic residuals, and combining trajectory correction with an intrinsic guidance objective, the method distills conditional information directly into network weights to optimize the inference trajectory. Notably, it eliminates the need for Classifier-Free Guidance, substantially reducing computational overhead. The approach achieves nearly threefold faster inference while preserving high timbre fidelity and significantly outperforms state-of-the-art baselines in speaker similarity.
This work addresses the challenge that large language models struggle to reliably self-verify their solutions in mathematical reasoning, and existing approaches suffer from high training costs and low efficiency. The authors propose a unified training framework that integrates problem solving and verification into a single generation process, introducing a Dynamic Reference Model Update (DRMU) mechanism combined with reward-based reinforcement learning for joint optimization. This approach significantly enhances self-verification performance, outperforming state-of-the-art methods across multiple mathematical benchmarks while reducing training time to only 51%–71% of prior approaches. The study also reveals the critical role of model scale in verification capability.
研究解决了全双工对话系统中区分打断和反馈的问题,通过创建ECHO基准测试,使用匹配对比方法评估上下文敏感的决策。
为解决全双工对话模型数据不足的问题,通过构建包含情景、行为、情感和声音事件的DuplexDrama合成对话数据集来提供支持。
为解决多轮代理任务中稀疏终端奖励导致的中间动作信用分配不精细问题,提出基于势能引导的策略优化方法PGPO,通过估计状态势能并传播跨轨迹信用。
Flow Matching in speech synthesis suffers from high inference latency and timbre leakage. This work proposes a unified guidance framework that, for the first time, jointly integrates data-level and model-level guidance. By leveraging heterogeneous data augmentation to disentangle linguistic content from acoustic residuals, and combining trajectory correction with an intrinsic guidance objective, the method distills conditional information directly into network weights to optimize the inference trajectory. Notably, it eliminates the need for Classifier-Free Guidance, substantially reducing computational overhead. The approach achieves nearly threefold faster inference while preserving high timbre fidelity and significantly outperforms state-of-the-art baselines in speaker similarity.
This work addresses the challenge that large language models struggle to reliably self-verify their solutions in mathematical reasoning, and existing approaches suffer from high training costs and low efficiency. The authors propose a unified training framework that integrates problem solving and verification into a single generation process, introducing a Dynamic Reference Model Update (DRMU) mechanism combined with reward-based reinforcement learning for joint optimization. This approach significantly enhances self-verification performance, outperforming state-of-the-art methods across multiple mathematical benchmarks while reducing training time to only 51%–71% of prior approaches. The study also reveals the critical role of model scale in verification capability.