Can LLMs Discover Scientific Laws in Real and Parallel Worlds?

📅 2026-09-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过引入SCILAWS-BENCH基准,评估大型语言模型在真实和合成环境中发现科学定律的能力,涵盖多个学科的真实数据和问题。
📝 Abstract
Scientific equation discovery has long been central to scientific progress, proceeding through iterative cycles of hypothesis generation, observational testing, and refinement under scientific constraints. As LLM capabilities advance and their role in AI for Science expands, it remains an open problem whether they can genuinely discover scientific laws and how this ability should be evaluated. Existing evaluations, however, often either simplify discovery through synthetic settings or reuse published targets that may already be familiar to LLMs. We therefore introduce SCILAWS-BENCH, a benchmark for scientific law discovery built from published research and real scientific data. It comprises 118 problems drawn from 381 scientific papers, covering 291 candidate laws and roughly 8M real data points across six scientific disciplines. Each problem is instantiated in two complementary settings: (1) SCILAWS-REAL asks models to propose laws from fixed real observations and evaluates held-out predictive fit and scientific validity derived from the source literature, and (2) SCILAWS-PARALLEL asks models to actively query residual-calibrated worlds and recover synthesized hidden laws derived from published forms. This two-setting task design preserves each problem's scientific context while separately evaluating fixed-record law discovery and active recovery of a newly synthesized hidden law. We find that predictive fit can diverge from scientific validity, memorization shapes whether models reproduce or move beyond published formulas, and our best-of-N study reveals a selection bottleneck. Our work provides a paper-grounded benchmark and new empirical perspectives for evaluating AI for scientific discovery. Project page: https://yiyihum.github.io/SciLaws-Bench
Problem

Research questions and friction points this paper is trying to address.

Scientific Law Discovery
Large Language Models
Benchmark Evaluation
Real Scientific Data
Active Recovery
Innovation

Methods, ideas, or system contributions that make the work stand out.

SCILAWS-BENCH
scientific law discovery
large language models
real-world data
parallel worlds
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yiming Huang
University of California San Diego
Ziche Liu
Ziche Liu
The Chinese University of Hong Kong, Shenzhen
NLPLLM & Instruction TuningWorld Model
Z
Zhuohang Wu
University of California, Irvine
Yiqian Wang
Yiqian Wang
Cushing Academy
J
Junxia Cui
University of California San Diego
X
Xinkai Zou
University of California San Diego
L
Linjun Mao
University of California San Diego
N
Nan Huang
University of California San Diego
N
Naicheng Yu
University of California San Diego
Kaijie Zhu
Kaijie Zhu
University of California, Santa Barbara
Yue Ma
Yue Ma
Bytedance
NLPDialogue SystemLLM
K
Kun Zhou
University of California San Diego
Letian Peng
Letian Peng
UC San Diego
MLNLPLLM
Jingbo Shang
Jingbo Shang
Associate Professor, UC San Diego
Natural Language ProcessingData MiningDeep LearningInformation ExtractionWeak Supervision