OdinEval: A Reproducible Benchmark for LLM-Based Program Repair in the Odin Programming Language

📅 2026-08-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为了解决Odin语言中程序修复基准不足的问题,研究构建了OdinEval基准,通过文档化的缺陷和特定测试来评估六种语言模型在168个实例上的修复性能。
📝 Abstract
Repository-level repair benchmarks still center on a few mainstream languages, leaving systems languages such as Odin largely untested. We present OdinEval, a reproducible benchmark built from documented defects in public Odin repositories. Each instance binds an issue to base and fix commits, a gold patch, an issue-specific regression test, a historical toolchain, and execution records. Admission requires the test to fail on the base revision and pass after the gold fix. When no usable developer test exists, a black-box test is reviewed independently by three instances of the same model, executed in both historical states, and revised from recorded feedback under a versioned Test Writing Skill. We evaluate six language models on 168 filtered instances under one shared protocol. Kimi-K3 records the highest Resolved score at 66.7%, while Qwen3.8-Max has the highest Repro score at 96.4%. The release includes frozen data, source archives, containers, validators, model patches, and audit manifests.
Problem

Research questions and friction points this paper is trying to address.

Odin Programming Language
Program Repair
Benchmark
Reproducibility
Systems Languages
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reproducible Benchmark
Program Repair
Odin Programming Language
Black-Box Test
Test Writing Skill
B
Bang Xie
Shanghai Jiao Tong University, Shanghai, China
Hao Liu
Hao Liu
University of Electronic Science and Technology of China
RISstacked intelligent metasurfaceDRL
Zhiyuan Peng
Zhiyuan Peng
Shanghai Jiao Tong University
Code GenerationBlockchain
Xin Yin
Xin Yin
Zhejiang University
LLM4CodeAgent
S
Senjian Zhang
Shanghai Jiao Tong University, Shanghai, China
Y
Yuan Luo
Shanghai Jiao Tong University, Shanghai, China
C
Chenhao Ying
Shanghai Jiao Tong University, Shanghai, China
Haiming Jin
Haiming Jin
Shanghai Jiao Tong University
wireless sensingreinforcement learning
W
Wei Chen
Shanghai Jiao Tong University, Shanghai, China
S
Shaocong Long
Shanghai Jiao Tong University, Shanghai, China
Z
Zhenyu Shi
Shanghai Jiao Tong University, Shanghai, China