Confidence-Gated Transductive Test Generation for Code Reranking

📅 2026-09-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决高质量测试用例难以获取的问题,提出了一种基于置信度门控的归纳-演绎测试生成方法(CoTT),有效提升了输出可靠性并降低了计算成本。
📝 Abstract
Test case synthesis is crucial for evaluating and ranking programs generated by large language models (LLMs). However, constructing high-quality test cases remains challenging because reliable expected outputs are often difficult to obtain. We propose Confidence-Gated Transductive Test Generation (CoTT), which first uses an efficient inductive procedure and invokes transductive generation only when inductive confidence is low. This adaptive design improves output reliability while allocating extra computation only when needed. On code reranking benchmarks, CoTT outperforms prior baselines across the reported metrics while reducing cost relative to applying transductive generation to every input. These results show that confidence-based allocation of test-time computation provides a favorable efficiency-effectiveness trade-off with a single efficient LLM.
Problem

Research questions and friction points this paper is trying to address.

test case synthesis
large language models
expected outputs
Innovation

Methods, ideas, or system contributions that make the work stand out.

Confidence-Gated
Transductive Test Generation
Code Reranking
Adaptive Design
🔎 Similar Papers
No similar papers found.
S
Sungjae Lee
Department of Computer Science and Engineering, POSTECH, South Korea
Y
Youngsik Yoon
Department of Computer Science and Engineering, POSTECH, South Korea
S
Seockbean Song
Graduate School of Artificial Intelligence, POSTECH, South Korea
Siwei Wang
Siwei Wang
National University of Defense Technology
Large-graph studymulti-view fusionmulti-view clustering
W
Wei Chen
Microsoft Research Asia, China
Jungseul Ok
Jungseul Ok
Associate Professor, CSE/AI, POSTECH
Reinforcement LearningMachine Learning