XREPOTEST: Benchmarking Multilingual Repository-Level Unit Test Generation for Large Language Models

📅 2026-08-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决大语言模型在多语言仓库级别单元测试生成中的实际应用问题,提出XREPOTEST基准,使用多种上下文增强策略评估14种先进模型的性能。
📝 Abstract
Large language models (LLMs) have shown promise for automated unit test generation, but existing evaluations largely rely on standalone settings and a narrow set of programming languages, overestimating real-world readiness. We introduce XREPOTEST, a multilingual repository-level benchmark for unit test generation spanning five underexplored languages: Rust, Go, Julia, PHP, and Ruby. XREPOTEST evaluates tests under realistic repository constraints using a containerized execution framework and multiple context augmentation strategies, including file-level, LSP-based, and retrieval-based context. Beyond standard metrics such as test pass rate and coverage, we propose Invocation Rate (IR) to assess whether generated tests meaningfully exercise the intended functionality. Experiments with 14 state-of-the-art LLMs, including Claude 4.5, GPT-5.2, DeepSeek V4-Pro, and Qwen families, reveal a substantial gap between standalone and repository-level performance, as well as trade-offs between richer context and test reliability. Overall, XREPOTEST provides a challenging and informative benchmark to advance scalable and robust unit test generation in realistic software environments. The dataset and code are publicly available at: https://github.com/solis-team/XRepoTest
Problem

Research questions and friction points this paper is trying to address.

Large language models
Unit test generation
Multilingual benchmark
Repository-level evaluation
Realistic software environments
Innovation

Methods, ideas, or system contributions that make the work stand out.

multilingual repository-level benchmark
context augmentation strategies
Invocation Rate (IR)
💼 Related Jobs
No related jobs found.
D
Dung Le Quang
Hanoi University of Science and Technology, Viet Nam
D
Dong Cao Van
Hanoi University of Science and Technology, Viet Nam
Nam Le Hai
Nam Le Hai
Hanoi University of Science and Technology
NLPAI4CodeAI4SEContinual learning
Linh Ngo Van
Linh Ngo Van
Hanoi University of Science and Technology
Machine LearningData MiningNatural Language Processing
A
Anh M. T. Bui
Hanoi University of Science and Technology, Viet Nam
P
Phuong T. Nguyen
University of L'Aquila, Italy