When the API Speaks the Wrong Language: Revisiting Post-Training for Multilingual Tool Use

📅 2026-08-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of argument language mismatch (ALM) in multilingual API calling with large language models, where semantically correct calls often fail due to inconsistent parameter languages. The authors systematically evaluate post-training strategies and propose a hybrid approach combining supervised fine-tuning (SFT) with reinforcement learning using structured, parameter-aware rewards—specifically GRPO. Their findings reveal that SFT alone substantially improves both argument language consistency and function-calling accuracy, achieving performance on par with or even surpassing more complex reinforcement learning methods. While GRPO provides modest gains in generalization and multi-objective trade-offs, the results underscore SFT as a strong baseline for multilingual tool usage, highlighting its underappreciated efficacy in this setting.
📝 Abstract
The reliability of Large Language Models (LLMs) for API calling degrades in multilingual settings. A common failure occurs when a model selects the correct tool but generates argument values in an inconsistent language, which we term Argument Language Mismatch (ALM). Although semantically correct, such outputs are operationally invalid and not captured by standard API-calling metrics. We revisit post-training strategies for mitigating ALM and find that, in our benchmark, supervised fine-tuning (SFT) provides a strong baseline, substantially improving argument language consistency and end-to-end function call accuracy. Under consistent model selection, SFT achieves performance comparable to, and sometimes exceeding more complex reinforcement learning (RL) approaches. We further examine whether RL with structured, argument-aware rewards offers additional benefits. While methods such as Group Relative Policy Optimization (GRPO) can improve language consistency and better preserve general reasoning ability, these gains are incremental and most pronounced in generalization and multi-objective trade-offs. Overall, our results suggest that much of the performance in multilingual API grounding can be achieved through careful supervised training, with RL providing targeted rather than fundamental improvements.
Problem

Research questions and friction points this paper is trying to address.

multilingual API calling
Argument Language Mismatch
language consistency
tool use
LLM reliability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Argument Language Mismatch
Supervised Fine-Tuning
Multilingual API Calling
Reinforcement Learning
Post-Training
💼 Related Jobs
No related jobs found.
S
Siddharth Chauhan
T
Thomas Butler
Abhishek Singhania
Abhishek Singhania
Amazon
Natural Language Processing
P
Pankaj Porwal
Honey Gupta
Honey Gupta
Amazon