Agentic Auto-Research is Fuzz Testing

📅 2026-08-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of sparse feedback faced by autonomous scientific agents when exploring vast experimental spaces, where conventional generate-and-rank paradigms struggle to guide effective exploration. Drawing an analogy between autonomous research and gray-box fuzz testing, the study proposes a coverage-guided dense feedback mechanism that leverages intermediate experimental executions to obtain low-cost, high-frequency signals of cognitive progress, dynamically steering subsequent interventions. The approach integrates a feedback-driven search strategy with a safeguarded scientific validation pipeline to mitigate false discoveries arising from adaptive reuse. Experimental results demonstrate that the proposed mechanism effectively predicts genuine scientific progress, substantially improves the cost-efficiency of exploration, and reduces the rate of false discoveries.
📝 Abstract
Autonomous research agents can generate experiments faster than researchers can validate them. Researchers have responded by scaling the proposer and ranking more samples with a learned judge or human reviewers. We argue that this *generate-and-rank* paradigm misses the problem of sparse feedback. Within a declared research problem, an agent follows the control loop of a greybox fuzzer: it proposes a candidate, executes it, observes feedback, and chooses what to try next. A fuzzer rarely finds a bug, but coverage makes partial progress observable on every execution. Fuzzers then use that signal to mutate inputs and allocate effort, rather than only to rank completed runs. Auto-research needs the same two capabilities. First, each experiment should expose a cheap, dense signal of epistemic progress before final scientific validation is available. Second, that signal should determine the next intervention so that the agent searches rather than repeatedly samples. Because the optimized progress signal is guidance rather than a verdict, final validation must still decide what counts as a discovery using evidence protected from adaptive reuse. We propose controlled tests of whether candidate signals predict validated progress, whether feedback-directed search yields more validated discoveries per unit cost than repeated sampling, and whether protected validation reduces false discoveries. Feedback architecture, not only generation, is a central bottleneck in auto-research.
Problem

Research questions and friction points this paper is trying to address.

Agentic Auto-Research
Sparse Feedback
Fuzz Testing
Epistemic Progress
Feedback Architecture
Innovation

Methods, ideas, or system contributions that make the work stand out.

fuzz testing
autonomous research agents
feedback-directed search
epistemic progress signal
protected validation
💼 Related Jobs
No related jobs found.
Y
Yifeng He
University of California, Davis
J
Jicheng Wang
University of California, Davis
Y
Yinzhe Zhao
Zhejiang University
J
Jiachen Liu
ARA Lab
H
Hao Chen
The University of Hong Kong