SFT or RL for Tool-Calling Agents? A Controlled Study Across Data, Method, and Scale
研究通过控制实验,比较了SFT与LoRA、RL(GRPO)及两者结合的方法在不同规模Qwen3模型上工具调用性能的影响,发现SFT与LoRA在多数情况下表现最佳。
研究通过控制实验,比较了SFT与LoRA、RL(GRPO)及两者结合的方法在不同规模Qwen3模型上工具调用性能的影响,发现SFT与LoRA在多数情况下表现最佳。
研究通过引入Jaccard路由重叠和适配器梯度余弦相似性,诊断MoE+LoRA微调中的领域间竞争问题,并提出SpawnLoRA方法来动态增加门控子适配器以减少负面迁移。
研究针对客服中心实时对话中的主题匹配问题,通过比较正则表达式、句嵌入编码器和基于Gemini的LLM匹配器,发现轻量级LLM匹配器在使用自然语言描述时表现最佳。
This work addresses the challenge faced by non-technical users in enterprise environments who struggle to securely and compliantly access governed analytical data through natural language. Traditional text-to-SQL approaches fail to accommodate managed APIs that encapsulate complex business logic. To bridge this gap, the authors propose Analytic Agent—a large language model (LLM)-based agent system that translates natural language intents into secure invocations of governed analytical APIs through multi-step reasoning and policy-aware orchestration. The system integrates permission validation, query execution, and compliance-aware visualization, achieving the first deep integration of LLM agents with enterprise-grade governed APIs while ensuring auditability, consistency, and security. Evaluation on 90 real-world enterprise use cases demonstrates that the system accurately interprets user intent, executes compliant queries, and generates visual results, significantly enhancing self-service analytics capabilities for non-technical users.
Existing benchmarks for evaluating tool use in spoken-language agents lack realistic audio scenarios. This work proposes a dataset-agnostic framework that automatically converts existing text-based tool-use benchmarks into paired speech versions without requiring re-annotation, by leveraging text-to-speech synthesis, multi-speaker voice conversion, and environmental noise injection. The approach innovatively introduces a reference-free LLM-as-judge protocol coupled with ambiguity-aware stress testing, enabling verifiable evaluation under privacy-preserving conditions. Evaluations of seven models on spoken variants of Confetti and When2Call reveal strong dependencies of performance on both model architecture and task characteristics. An open-source Qwen3 judge model achieves over 80% agreement with commercial counterparts, demonstrating the framework’s validity and practical utility.
研究通过控制实验,比较了SFT与LoRA、RL(GRPO)及两者结合的方法在不同规模Qwen3模型上工具调用性能的影响,发现SFT与LoRA在多数情况下表现最佳。
研究通过引入Jaccard路由重叠和适配器梯度余弦相似性,诊断MoE+LoRA微调中的领域间竞争问题,并提出SpawnLoRA方法来动态增加门控子适配器以减少负面迁移。
研究针对客服中心实时对话中的主题匹配问题,通过比较正则表达式、句嵌入编码器和基于Gemini的LLM匹配器,发现轻量级LLM匹配器在使用自然语言描述时表现最佳。
This work addresses the challenge faced by non-technical users in enterprise environments who struggle to securely and compliantly access governed analytical data through natural language. Traditional text-to-SQL approaches fail to accommodate managed APIs that encapsulate complex business logic. To bridge this gap, the authors propose Analytic Agent—a large language model (LLM)-based agent system that translates natural language intents into secure invocations of governed analytical APIs through multi-step reasoning and policy-aware orchestration. The system integrates permission validation, query execution, and compliance-aware visualization, achieving the first deep integration of LLM agents with enterprise-grade governed APIs while ensuring auditability, consistency, and security. Evaluation on 90 real-world enterprise use cases demonstrates that the system accurately interprets user intent, executes compliant queries, and generates visual results, significantly enhancing self-service analytics capabilities for non-technical users.
Existing benchmarks for evaluating tool use in spoken-language agents lack realistic audio scenarios. This work proposes a dataset-agnostic framework that automatically converts existing text-based tool-use benchmarks into paired speech versions without requiring re-annotation, by leveraging text-to-speech synthesis, multi-speaker voice conversion, and environmental noise injection. The approach innovatively introduces a reference-free LLM-as-judge protocol coupled with ambiguity-aware stress testing, enabling verifiable evaluation under privacy-preserving conditions. Evaluations of seven models on spoken variants of Confetti and When2Call reveal strong dependencies of performance on both model architecture and task characteristics. An open-source Qwen3 judge model achieves over 80% agreement with commercial counterparts, demonstrating the framework’s validity and practical utility.