BENCHCOMPASS: From Scores to Signals for Training and Harness Decisions in Payment-Domain LLMs
为了解决支付领域大语言模型的应用问题,本文提出了BENCHCOMPASS基准测试,通过构建基于场景的任务和质量检查来评估模型处理支付规则、证据利用及鲁棒性。
为了解决支付领域大语言模型的应用问题,本文提出了BENCHCOMPASS基准测试,通过构建基于场景的任务和质量检查来评估模型处理支付规则、证据利用及鲁棒性。
本文提出MUSE-Bench,一个用于多模态时间序列预测的统一基准,包含多种类型上下文,评估不同预测方法的效果。
Reusable skills give agents transferable procedural knowledge, making scalable acquisition essential for extending agents beyond prior experience. Existing methods face two limitations: trajectory-based synthesis requires interactions with specific environments, while document-derived skills may lack executable evidence and verification. Source code offers a complementary path: it requires no prior agent experience yet provides executable evidence for grounding abstractions. We present Code2Skill, a fully automated pipeline that transforms selected code units into implementation-anchored records of atomic operations, composite workflows, and recurring patterns, then verifies each record through source-body-blind reconstruction and source-aware comparison. Applied to 19,769 popular, actively maintained GitHub repositories, Code2Skill produces CodeSkillBank, a grounded bank of 1,006,822 accepted records with workflow, boundary, provenance, and source-evidence metadata. Across 72 protocol-matched evaluations covering nine model settings and eight benchmarks, models augmented with retrieved CodeSkillBank skills improve by 11.7% on average over matched baselines and outperform them in 57 cases. Under a unified downstream interface, Code2Skill also outperforms trajectory-derived skill banks on all seven shared benchmarks, showing that repository-derived skills can provide useful procedural knowledge before agents accumulate sufficient interaction experience. Skills synthesized from tested AI-generated code achieve a 93.50% pass rate, compared with 93.00% for human-written code, providing initial evidence that the pipeline can expand with the growing volume of AI-generated software. Overall, Code2Skill transforms procedural knowledge embedded in repositories into grounded, verifiable, and transferable agent skills.
研究通过构建FinRiskAtlas基准,评估大型语言模型在金融风险审查中的特定操作执行与证据状态控制能力,以解决现有评估标准无法准确反映模型在专业工作流程中可靠性的不足。
本文提出LeakGauge方法,通过在响应中添加后缀来检测大语言模型处理外部上下文时的泄露风险,该方法在11个模型上表现出高稳定性与准确性。
为了解决支付领域大语言模型的应用问题,本文提出了BENCHCOMPASS基准测试,通过构建基于场景的任务和质量检查来评估模型处理支付规则、证据利用及鲁棒性。
本文提出MUSE-Bench,一个用于多模态时间序列预测的统一基准,包含多种类型上下文,评估不同预测方法的效果。
Reusable skills give agents transferable procedural knowledge, making scalable acquisition essential for extending agents beyond prior experience. Existing methods face two limitations: trajectory-based synthesis requires interactions with specific environments, while document-derived skills may lack executable evidence and verification. Source code offers a complementary path: it requires no prior agent experience yet provides executable evidence for grounding abstractions. We present Code2Skill, a fully automated pipeline that transforms selected code units into implementation-anchored records of atomic operations, composite workflows, and recurring patterns, then verifies each record through source-body-blind reconstruction and source-aware comparison. Applied to 19,769 popular, actively maintained GitHub repositories, Code2Skill produces CodeSkillBank, a grounded bank of 1,006,822 accepted records with workflow, boundary, provenance, and source-evidence metadata. Across 72 protocol-matched evaluations covering nine model settings and eight benchmarks, models augmented with retrieved CodeSkillBank skills improve by 11.7% on average over matched baselines and outperform them in 57 cases. Under a unified downstream interface, Code2Skill also outperforms trajectory-derived skill banks on all seven shared benchmarks, showing that repository-derived skills can provide useful procedural knowledge before agents accumulate sufficient interaction experience. Skills synthesized from tested AI-generated code achieve a 93.50% pass rate, compared with 93.00% for human-written code, providing initial evidence that the pipeline can expand with the growing volume of AI-generated software. Overall, Code2Skill transforms procedural knowledge embedded in repositories into grounded, verifiable, and transferable agent skills.
研究通过构建FinRiskAtlas基准,评估大型语言模型在金融风险审查中的特定操作执行与证据状态控制能力,以解决现有评估标准无法准确反映模型在专业工作流程中可靠性的不足。
本文提出LeakGauge方法,通过在响应中添加后缀来检测大语言模型处理外部上下文时的泄露风险,该方法在11个模型上表现出高稳定性与准确性。