VERPO: Verified Evidence Regularized Policy Optimization
为解决语言模型后训练中序列级优势无法指导具体决策的问题,VERPO通过引入验证证据调节策略优化方法,提高任务成功率。
为解决语言模型后训练中序列级优势无法指导具体决策的问题,VERPO通过引入验证证据调节策略优化方法,提高任务成功率。
FireRedAudio通过解耦连续表示法解决音频理解与生成的统一问题,使用9B参数LLM实现长上下文建模和细粒度声学细节保留。
针对大型语言模型在动态环境中知识管理问题,提出双层代理记忆框架,通过成本感知的写入路由和周期性参数整合,有效管理知识生命周期。
本文提出OneModel框架,通过统一多场景排序解决平台级推荐系统中用户行为跨流连续性问题,采用共享事件序列和情景感知信息调节等方法增强用户表示。
This work addresses the limitations of generic Retrieval-Augmented Generation (RAG) approaches in enterprise data analytics, which suffer from low retrieval accuracy (Hit@10 of only 19.1%) and misinterpretation of metrics due to semantic gaps, entity ambiguity, schema drift, and misalignment between data assets and their usage contexts. To overcome these challenges, the authors propose a dual-layer architecture: the lower layer constructs a three-tier, dual-purpose knowledge base integrating a knowledge graph with 2,859 nodes and eight-segment scenario annotations; the upper layer introduces a graph-guided retriever (GGR) and a scenario-aware ranker (SAR), enhanced by negative knowledge augmentation and a lightweight closed-loop hot-reload mechanism enabling daily knowledge updates. Evaluated on two hundred-question benchmarks, the approach achieves a Hit@10 of 96.6%, increases knowledge coverage to 77%, and maintains end-to-end latency between 4.84 and 5.33 seconds.
为解决语言模型后训练中序列级优势无法指导具体决策的问题,VERPO通过引入验证证据调节策略优化方法,提高任务成功率。
FireRedAudio通过解耦连续表示法解决音频理解与生成的统一问题,使用9B参数LLM实现长上下文建模和细粒度声学细节保留。
针对大型语言模型在动态环境中知识管理问题,提出双层代理记忆框架,通过成本感知的写入路由和周期性参数整合,有效管理知识生命周期。
本文提出OneModel框架,通过统一多场景排序解决平台级推荐系统中用户行为跨流连续性问题,采用共享事件序列和情景感知信息调节等方法增强用户表示。
This work addresses the limitations of generic Retrieval-Augmented Generation (RAG) approaches in enterprise data analytics, which suffer from low retrieval accuracy (Hit@10 of only 19.1%) and misinterpretation of metrics due to semantic gaps, entity ambiguity, schema drift, and misalignment between data assets and their usage contexts. To overcome these challenges, the authors propose a dual-layer architecture: the lower layer constructs a three-tier, dual-purpose knowledge base integrating a knowledge graph with 2,859 nodes and eight-segment scenario annotations; the upper layer introduces a graph-guided retriever (GGR) and a scenario-aware ranker (SAR), enhanced by negative knowledge augmentation and a lightweight closed-loop hot-reload mechanism enabling daily knowledge updates. Evaluated on two hundred-question benchmarks, the approach achieves a Hit@10 of 96.6%, increases knowledge coverage to 77%, and maintains end-to-end latency between 4.84 and 5.33 seconds.