FrontierScience: Evaluating AI's Ability to Perform Expert-Level Scientific Tasks
Existing scientific evaluation benchmarks predominantly rely on multiple-choice questions or established knowledge, making them inadequate for assessing expert-level reasoning capabilities of AI systems in cutting-edge scientific tasks. To address this gap, this work introduces FrontierScience, a novel benchmark comprising two tracks: one featuring International Olympiad-level problems and the other consisting of open-ended, doctoral-level research subtasks spanning frontier topics in physics, chemistry, and biology—such as quantum electrodynamics and synthetic organic chemistry. For the first time, the benchmark incorporates high-difficulty problems authored by Olympiad gold medalists and active scientists, and it employs a process-oriented, fine-grained scoring framework that moves beyond conventional answer-only evaluation. Comprising hundreds of high-quality questions—including 160 open-source “gold” items—the benchmark effectively discriminates among state-of-the-art models in advanced scientific reasoning.