Salesforce Koa: An Enterprise Language Model for Agentic Tool Use

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过使用Group Relative Policy Optimization对Nemotron-3-Super-120B模型进行后训练,开发了Salesforce Koa语言模型以改善工具使用和代理能力。
📝 Abstract
We present Salesforce Koa, an enterprise language model built by post-training the open-weight Nemotron-3-Super-120B foundation model with reinforcement learning using Group Relative Policy Optimization (GRPO). Salesforce Koa is trained on public and synthetically generated data, with no customer data, to improve tool use and agentic capabilities while preserving strong general-purpose performance. Its distinctive component is a simulation-to-reward pipeline that expands workflow specifications into persona-conditioned multi-turn tasks with task-resolution rewards grounded in successful tool use for data-dependent requests. For enterprise domains, these specifications are written in Agent Script, Salesforce's declarative language for building Agentforce agents; for public tool-use domains, we synthesize the workflow structure directly. The same simulation and grounded-reward machinery drives GRPO across both. Across public tool-use, agentic-reasoning, and enterprise Customer Relationship Management (CRM) benchmarks, Salesforce Koa improves over its open-weight base, with the clearest gains on multi-turn tool use, and surpasses a strong proprietary baseline while remaining below the strongest frontier models. These results show that specification-driven reinforcement learning is a practical path to specializing open-weight foundation models for enterprise agentic tasks.
Problem

Research questions and friction points this paper is trying to address.

enterprise language model
tool use
agentic capabilities
general-purpose performance
reinforcement learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Group Relative Policy Optimization
Simulation-to-reward Pipeline
Persona-conditioned Multi-turn Tasks
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Z
Zixiang Chen
Salesforce Agentforce & AI Research
Sufeng Niu
Sufeng Niu
Salesforce Agentforce & AI Research
Y
Yingchi Liu
Salesforce Agentforce & AI Research
Wenting Zhao
Wenting Zhao
Salesforce Research
Program SynthesisNatural Language ProcessingData Mining
A
Akshara Prabhakar
Salesforce Agentforce & AI Research
S
Shubham Mehrotra
Salesforce Agentforce & AI Research
B
Bin Bi
Salesforce Agentforce & AI Research
Z
Zhujun Lan
Salesforce Agentforce & AI Research
K
Katherine Tan
Salesforce Agentforce & AI Research
M
Mohammad Ramezanali
Salesforce Agentforce & AI Research
T
Tulika Manoj Awalgaonkar
Salesforce Agentforce & AI Research
Monojit Banerjee
Monojit Banerjee
Georgia Institute of Technology
AIMachine LearningDevOps
Jielin Qiu
Jielin Qiu
Research Scientist, Salesforce AI Research
Machine LearningMultimodal Learning
S
Shiva Kumar Pentyala
Salesforce Agentforce & AI Research
Z
Zhepeng Cen
Salesforce Agentforce & AI Research
A
Anupam Tripathi
Salesforce Agentforce & AI Research
A
Ali Ziaei
Salesforce Agentforce & AI Research
R
Regunathan Radhakrishnan
Salesforce Agentforce & AI Research
D
Darvish Lee Shadravan
Salesforce Agentforce & AI Research
Shelby Heinecke
Shelby Heinecke
Salesforce Research
Artificial IntelligenceAI AgentsLLM AgentsMulti-Agent SystemsRecommendation Systems
Sitaram Asur
Sitaram Asur
Director, Salesforce (formerly in HP Labs)
Machine LearningNLPSocial Networks
Silvio Savarese
Silvio Savarese
Associate Professor of Computer Science at Stanford University
Computer vision
J
James Zhu
Salesforce Agentforce & AI Research
P
Phil Mui
Salesforce Agentforce & AI Research
Huan Wang
Huan Wang
Salesforce Research
Machine LearningArtificial IntelligenceComputer Vision