Auditing and Decomposing Feedback-Driven Evolution in LLM Test Generation under the Oracle Problem

📅 2026-08-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究审计了反馈驱动测试生成中的问题,使用多种任务和模型对比独立重采样与变异进化方法,提出审计及安慰剂协议以分离验证器伪像、交互框架和真实反馈效果。
📝 Abstract
Execution feedback is often treated as a self-verifying signal for improving LLM-generated tests. However, when generated inputs are executed on a single accepted program and its outputs are used as ground truth, invalid or underspecified inputs can create spurious fault detections and apparent evolutionary gains. We audit this failure mode in feedback-driven test generation using 142 development tasks, 114 locked external tasks, and 138 held-out tasks, with two code models, three seeds, and fault-cross-fitted real submissions. On external inputs for which three accepted implementations agree, generated outputs match the panel on only 27.79% and 50.12% of cases. A single-reference oracle inflates the measured gain from evolution by 9.46-14.85 percentage points; after auditing, equal-budget independent resampling outperforms mutation-based evolution by 6.01-18.83 points. We further compare a genuine three-round feedback loop with a density-matched placebo. External Real-Placebo differences are +0.13 and -0.50 points, while held-out differences are +1.99 and +0.28 points and do not provide robust evidence of fine-grained feedback benefit. A blinded semantic audit by two software engineering doctoral students classifies 94.41% of panel-disconfirmed inputs as invalid but 3.60% as valid, showing that panel disagreement is informative but not semantic proof. We propose an audit-and-placebo protocol that separates verifier artifacts, interaction scaffolding, and grounded feedback credit in evaluations of self-evolving test generators.
Problem

Research questions and friction points this paper is trying to address.

execution feedback
LLM test generation
oracle problem
spurious fault detections
evolutionary gains
Innovation

Methods, ideas, or system contributions that make the work stand out.

audit-and-placebo protocol
feedback-driven test generation
independent resampling
💼 Related Jobs
No related jobs found.
Y
Yunhao Liang
Chengdu Institute of Computer Applications, Chinese Academy of Sciences, China and University of Chinese Academy of Sciences, China
C
Chengguang Gan
Independent Researcher, Japan
R
Ruixuan Ying
Institute of Multidisciplinary Research for Advanced Materials (IMRAM), Tohoku University, Japan
H
Hanjun Wei
University of Chinese Academy of Sciences, China
Zhe Cui
Zhe Cui
Beijing University of Posts and Telecommunications
fingerprint
S
Shiwen Ni
Artificial Intelligence Research Institute, Shenzhen University of Advanced Technology, China