SFT or RL for Tool-Calling Agents? A Controlled Study Across Data, Method, and Scale

📅 2026-09-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过控制实验,比较了SFT与LoRA、RL(GRPO)及两者结合的方法在不同规模Qwen3模型上工具调用性能的影响,发现SFT与LoRA在多数情况下表现最佳。
📝 Abstract
Limited controlled evidence exists on how training data, adaptation method, and model scale jointly affect tool-calling performance in language-model agents. We evaluate supervised fine-tuning (SFT) with LoRA, reinforcement learning (RL) via Group Relative Policy Optimization (GRPO), and SFT followed by GRPO across six Qwen3 models from 0.6B to 32B parameters, covering both in-distribution performance and cross-dataset transfer. SFT with LoRA is the strongest in-distribution method throughout the 0.6B-32B range and best in 15 out of 18 experimental settings. On cross-dataset transfer, the methods are closer: GRPO wins 29 out of 54 settings where training and test datasets differ, but its margin over SFT averages under one point, and SFT->GRPO is rarely strongest in either comparison. Dataset mixing gives consistently strong transfer while staying close to specialized in-distribution training, regardless of method. Additional analysis further confirms that LoRA outperforms full-parameter fine-tuning, demonstrating that LoRA better preserves pretrained agentic behavior.
Problem

Research questions and friction points this paper is trying to address.

training data
adaptation method
model scale
tool-calling performance
language-model agents
Innovation

Methods, ideas, or system contributions that make the work stand out.

Supervised Fine-Tuning (SFT)
LoRA
Reinforcement Learning (RL)
Cross-Dataset Transfer
Group Relative Policy Optimization (GRPO)
🔎 Similar Papers