ASI-Bench: At the Dawn of Artificial Superintelligence

📅 2026-08-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文介绍了ASI-Bench,旨在评估AI系统在探索未知、自主执行科学研究方面的能力,通过逐步减少人类指导来测试AI独立完成科研任务的水平。
📝 Abstract
Artificial superintelligence (ASI) requires AI to move beyond mastering existing knowledge toward exploring the unknown, creating new knowledge, and turning new ideas into verifiable results. However, the capabilities of today's AI systems are still largely built on learning, compressing, and applying existing human knowledge. Accordingly, existing benchmarks primarily test whether AI can produce correct answers based on learned knowledge, or whether it can complete tasks under extensive human guidance. We therefore introduce ASI-Bench, the first benchmark to jointly evaluate AI systems' capabilities of innovative exploration and autonomous scientific execution across general research domains, and the first to progressively withdraw human methodological guidance within the same research project to test how far AI can proceed on its own. Built by over 40 experts with the cost of 31,000+ human hours, ASI-Bench contains 60 project-level research tasks across 11 scientific domains and progressively reduces methodological guidance to test whether AI can independently select methods, conduct research, and produce verifiable results. All tasks undergo expert review, AI-assisted auditing, sandbox execution, and scorer validation. Across 18 state-of-the-art agent--model configurations, the average score drops from 50.91 with full methodological guidance to 29.10 with only the method specified and 26.62 when agents must determine the method themselves. This sharp decline shows that current systems remain heavily dependent on human guidance and are still far from autonomously conducting end-to-end, project-level scientific research. ASI-Bench is open to the world. We invite researchers and builders everywhere to contribute new tasks, challenge the limits of today's AI, and help accelerate humanity's collective path toward artificial superintelligence at https://asibench.apexin.ai/submit.
Problem

Research questions and friction points this paper is trying to address.

Artificial Superintelligence
Innovative Exploration
Autonomous Scientific Execution
Methodological Guidance
Innovation

Methods, ideas, or system contributions that make the work stand out.

ASI-Bench
Innovative Exploration
Autonomous Scientific Execution
Methodological Guidance Reduction
End-to-End Research
🔎 Similar Papers
Junwei Zhou
Junwei Zhou
Dartmouth College
Computer Vision
Zhen Sun
Zhen Sun
DSA Thrust, HKUST(GZ)
LLM security
B
Binyu Li
Tsinghua University
J
Jiangyu Zhou
Tsinghua University
Y
Yuexi Pan
Tsinghua University
H
Hengyu Wang
Tsinghua University
H
Honghe Ren
Tsinghua University
X
Xiaohan Jia
Tsinghua University
X
Xueyang Zhou
Tsinghua University
X
Xiaoyu Cao
Tsinghua University
Yongchao Chen
Yongchao Chen
Harvard University, Massachusetts Institute of Technology
Robot PlanningFoundation ModelsFormal MethodsMechanicsAI for Science
Y
Yuanning Feng
Tsinghua University
Junhao Wu
Junhao Wu
Towson university
Computer VisionCryo emMedical image
C
Cheng Zhang
Harvard University
S
Sijia Chen
Flatiron Institute
H
Haoyu Xue
Tsinghua University
C
Chengsong You
Tsinghua University
H
Huan Wang
Tsinghua University
K
Koutian Wu
Harvard University
P
Peigan Gao
University of Science and Technology of China
J
Jiakun Wu
Tsinghua University
Wenzhe Li
Wenzhe Li
Princeton University
E
Ergan Shang
Carnegie Mellon University
Qingyuan Zheng
Qingyuan Zheng
Tsinghua University
J
Jingjing Zhou
Tsinghua University