MathAdv: What Theorem Provers Know, Reason, Formalize, and Generalize

📅 2026-08-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该研究通过引入MathAdv基准,解决现有定理证明器评估标准的局限性问题,采用多任务测试方法全面评估模型在数学推理、知识和鲁棒性方面的能力。
📝 Abstract
Formal theorem proving enables machine-verifiable evaluation of mathematical reasoning, yet existing benchmarks often emphasize aggregate proof accuracy, concentrate on a narrow range of mathematics, and provide limited evidence of robustness to equivalent reformulations. We introduce MathAdv, a diagnostic benchmark spanning 13 domains across undergraduate- and graduate-level mathematics. Alongside Lean 4 theorem proving, MathAdv provides up to three auxiliary tasks: multiple-choice questions that probe mathematical knowledge, fill-in-the-blank problems that isolate informal reasoning, and expert-crafted transformations that test robustness to problem presentation. Our evaluation of contemporary theorem provers yields four findings: formalization remains a major bottleneck; performance varies substantially across mathematical domains; natural-language guidance helps general-purpose LLMs but can hinder proof-specialized models; and mathematically equivalent reformulations expose substantial robustness limitations. Together, these results show how component-wise evaluation can reveal model capabilities and failure modes that aggregate theorem-proving accuracy obscures. The dataset and evaluation scripts are available at https://github.com/margotyjx/MathAdv.git.
Problem

Research questions and friction points this paper is trying to address.

Theorem Proving
Benchmark
Mathematical Reasoning
Robustness
Innovation

Methods, ideas, or system contributions that make the work stand out.

diagnostic benchmark
theorem proving
mathematical reasoning
robustness to reformulations
natural-language guidance
🔎 Similar Papers
Jiaxin Yuan
Jiaxin Yuan
Ph.D. candidate, University of Maryland
Stochastic differential equationmolecular dynamicscausal inference
C
Connor Martinez Lockhart
University of Maryland, College Park
Xiaoyu Liu
Xiaoyu Liu
University of Maryland, College Park
Jiaqi Wang
Jiaqi Wang
Harbin Institute of Technology Shenzhen & Pengcheng Laboratory, Computer Science
Spiking Neural NetworkBrain DecodingSpeechBrain Computer Interface
Chenghao Deng
Chenghao Deng
University of Maryland, College Park
machine learninglarge language model
X
Xiayimei Han
University of Maryland, College Park
V
Vlasios Mastrantonis
Cornell University
D
Dmitrii Gudin
University of Maryland, College Park
S
Shaopeng Zhu
Independent Researcher
A
Abdirisak Abdullahi Mohamed
University of Maryland, College Park
B
Bilal Hamdi Aytekin
University of Maryland, College Park
J
Jiewen Lang
Independent Researcher
Z
Zezheng Song
University of Maryland, College Park
Furong Huang
Furong Huang
Associate Professor of Computer Science, University of Maryland
Trustworthy AI/MLReinforcement LearningGenerative AI