HypoForge: A Self-Improving Multi-Agent Framework for Automated Hypothesis Generation and Testing via Scientific Skill Learning

📅 2026-08-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
HypoForge通过学习科学技能,采用多代理框架解决自动假设生成和测试问题,利用阶段特定的监督信号实现持续改进。
📝 Abstract
Large language models (LLMs) have enabled AI scientist systems to automate scientific discovery, yet existing approaches most rely on static prompting or fixed workflows and fail to accumulate experience for continual improvement. We propose HypoForge, an experience-guided multi-agent framework that learns reusable scientific skills for automated hypothesis generation and hypothesis testing. HypoForge is built on the observation that these two stages involve different supervision signals. For hypothesis generation, where explicit feedback is unavailable, HypoForge adopts an adversarial generator--discriminator mechanism to improve reasoning through comparative critique. For hypothesis testing, where empirical feedback is available, HypoForge learns testing skills from execution outcomes and ground-truth results. By matching skill learning strategies with stage-specific supervision, HypoForge enables continual improvement without fine-tuning foundation models. Experiments on hypothesis generation and testing benchmarks show that HypoForge consistently outperforms existing AI scientist frameworks and skill-level variants. Further analysis demonstrates the effectiveness of the proposed stage-specific skill learning paradigms.
Problem

Research questions and friction points this paper is trying to address.

Large language models
Automated scientific discovery
Continual improvement
Hypothesis generation
Hypothesis testing
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-Improving Multi-Agent Framework
Scientific Skill Learning
Adversarial Generator--Discriminator Mechanism
Stage-Specific Supervision
💼 Related Jobs
No related jobs found.