Science sandboxes measure the scientific capability of AI agents

📅 2026-08-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究提出科学沙箱框架,通过实验、反馈和假设修正循环评估AI代理的科学能力,特别是在生物学模型中优化指标与理解系统规则之间的差异。
📝 Abstract
Scientific progress depends not only on finding solutions, but on learning the rules that explain why they work and using that understanding to design better experiments. We introduce science sandboxes, a framework for studying this capability in AI agents through repeated cycles of experimentation, feedback, and hypothesis revision. Science sandboxes invite an agent to query the natural world in different ways, ranging from "wet" physical experiments, to "damp" predictive models trained on empirical data, to "dry" invented rules. By establishing a common experimental loop and a protocol for evaluating agents within it, science sandboxes allow assessment of both quantitative performance on specific metrics and qualitative scientific reasoning, across a spectrum of empirical verifiability. Here, we instantiate this framework in two biological settings, models of regulatory genomics and protein fitness prediction, and examine the capabilities of frontier agents. Across these settings, we could see when agents successfully optimized a quantitative metric without understanding the rules underlying the system. In particular, their scientific reasoning deteriorated when they encountered systems whose rules fell outside familiar biological priors. By highlighting such failure modes, science sandboxes make the frontier of scientific capability measurable and provide a controlled setting in which to study and ultimately expand it.
Problem

Research questions and friction points this paper is trying to address.

AI agents
scientific capability
experimentation
hypothesis revision
biological settings
Innovation

Methods, ideas, or system contributions that make the work stand out.

science sandboxes
AI agents
scientific reasoning
experimental loop
biological settings
A
Arya S. Rao
The Broad Institute of MIT and Harvard; Cambridge, MA 02142, USA.
R
Rodrigo I. Castro
The Jackson Laboratory; Bar Harbor, ME 04609, USA.
S
Sager J. Gosai
Sutter Hill Ventures; Palo Alto, CA 94304, USA.
K
Kenneth B. Hsu
The Broad Institute of MIT and Harvard; Cambridge, MA 02142, USA.
Yasha Ektefaie
Yasha Ektefaie
Harvard Medical School, Biomedical Informatics
Machine LearningMedical InformaticsInfectious DiseaseBioinformatics
Shantanu Singh
Shantanu Singh
The Broad Institute of MIT and Harvard; Cambridge, MA 02142, USA.
S
Sangeeta N. Bhatia
The Broad Institute of MIT and Harvard; Cambridge, MA 02142, USA.; David H. Koch Institute for Integrative Cancer Research, Massachusetts Institute of Technology, Cambridge, MA 02139, USA.; Howard Hughes Medical Institute, Chevy Chase, MD 20815, USA.; The Wyss Institute for Biologically Inspired Engineering at Harvard University, Boston, MA 02115, USA.; Harvard-MIT Program in Health Sciences and Technology, Institute for Medical Engineering and Science, Massachusetts Institute of Technology, Cambridge, MA 0
S
Steven K. Reilly
Department of Genetics, Yale School of Medicine; New Haven, CT, USA.
Ryan Tewhey
Ryan Tewhey
The Jackson Laboratory; Bar Harbor, ME 04609, USA.
E
Eric S. Lander
The Broad Institute of MIT and Harvard; Cambridge, MA 02142, USA.; Department of Systems Biology, Harvard Medical School, Boston, MA, USA.; Department of Biology, Massachusetts Institute of Technology, Cambridge, MA, USA.
P
Pardis C. Sabeti
The Broad Institute of MIT and Harvard; Cambridge, MA 02142, USA.; Department of Immunology and Infectious Diseases, Harvard T.H. Chan School of Public Health, Boston, MA 02115, USA.; Department of Organismic and Evolutionary Biology, Harvard University, Cambridge, MA 02138, USA.