Evaluating Large Language Models for Code Translation: Effects of Prompt Language and Prompt Design

📅 2025-09-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study systematically evaluates large language models (LLMs) for cross-language code translation among C++, Java, Python, and C#, addressing a gap in empirical, multi-language, multi-prompt comparative analysis. Methodologically, it introduces the first comparison of English versus Arabic prompt languages, combines concise instructions with detailed specifications as two prompt styles, employs direction-aware paired evaluation, and quantifies performance using BLEU and CodeBLEU metrics, with TransCoder as a traditional baseline. Key contributions include: (1) demonstrating that all evaluated LLMs significantly outperform TransCoder; (2) revealing dual advantages of detailed prompts and English-language prompts—specifically, English prompts improve CodeBLEU by 13–15%; and (3) proposing evidence-based prompt engineering guidelines tailored to code translation, supporting practical software migration and cross-language interoperability. The findings provide rigorous, actionable insights for leveraging LLMs in real-world multilingual software engineering tasks.

Technology Category

Application Category

📝 Abstract
Large language models (LLMs) have shown promise for automated source-code translation, a capability critical to software migration, maintenance, and interoperability. Yet comparative evidence on how model choice, prompt design, and prompt language shape translation quality across multiple programming languages remains limited. This study conducts a systematic empirical assessment of state-of-the-art LLMs for code translation among C++, Java, Python, and C#, alongside a traditional baseline (TransCoder). Using BLEU and CodeBLEU, we quantify syntactic fidelity and structural correctness under two prompt styles (concise instruction and detailed specification) and two prompt languages (English and Arabic), with direction-aware evaluation across language pairs. Experiments show that detailed prompts deliver consistent gains across models and translation directions, and English prompts outperform Arabic by 13-15%. The top-performing model attains the highest CodeBLEU on challenging pairs such as Java to C# and Python to C++. Our evaluation shows that each LLM outperforms TransCoder across the benchmark. These results demonstrate the value of careful prompt engineering and prompt language choice, and provide practical guidance for software modernization and cross-language interoperability.
Problem

Research questions and friction points this paper is trying to address.

Evaluating LLMs for automated code translation across programming languages
Assessing effects of prompt design and language on translation quality
Comparing syntactic fidelity and structural correctness using BLEU metrics
Innovation

Methods, ideas, or system contributions that make the work stand out.

Detailed prompts improve translation quality
English prompts outperform Arabic by 13-15%
LLMs surpass traditional TransCoder baseline performance
🔎 Similar Papers
2024-03-252024 IEEE/ACM First International Conference on AI Foundation Models and Software Engineering (Forge) Conference Acronym:Citations: 22
💼 Related Jobs
No related jobs found.
A
Aamer Aljagthami
Department of Software Engineering, College of Computer Science and Engineering, University of Jeddah
M
Mohammed Banabila
Department of Software Engineering, College of Computer Science and Engineering, University of Jeddah
M
Musab Alshehri
Department of Software Engineering, College of Computer Science and Engineering, University of Jeddah
M
Mohammed Kabini
Department of Software Engineering, College of Computer Science and Engineering, University of Jeddah
M
Mohammad D. Alahmadi
Department of Software Engineering, College of Computer Science and Engineering, University of Jeddah