When Does AI for PDEs Yield Scientific Evidence?

📅 2026-08-23
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过扩展PDE仿真基准,首次评估AI模型输出对科学主张的支持程度,揭示了数值准确性和证据支持之间的不匹配问题。
📝 Abstract
Existing AI-for-PDE benchmarks primarily assess models in terms of predictive or approximation accuracy. In physics research, however, AI outputs often serve as evidence for scientific claims. These two objectives are not equivalent: the former measures an output's agreement with a reference target or satisfaction of governing constraints; the latter asks whether, given a specified object of study, scientific claim, assumptions, and evidence standard, the output provides sufficient evidence for that claim. To bridge this gap, we extend a widely used PDE-simulation benchmark and a comprehensive benchmark for PDE inverse problems to enable, for the first time in AI for PDEs, evaluation of whether and to what extent model outputs support specified scientific claims. Our results show that numerical accuracy and evidential support can rank models differently, explain when and why they do so, and reveal that existing benchmarks can favor methods whose outputs provide weaker support for the scientific claims of interest. Together, we formalize, empirically demonstrate, and explain this evaluation--use mismatch in AI for PDEs.
Problem

Research questions and friction points this paper is trying to address.

AI for PDEs
Scientific Evidence
Benchmarking
Prediction Accuracy
Evidential Support
Innovation

Methods, ideas, or system contributions that make the work stand out.

PDEs
scientific evidence
benchmark extension
evidential support