A Lean Dataset for International Math Olympiad: Small Steps towards Writing Math Proofs for Hard Problems

📅 2024-11-28
🏛️ arXiv.org
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
The scarcity of formalized proofs for International Mathematical Olympiad (IMO) problems hinders progress in AI-based mathematical reasoning. Method: We construct the first publicly available Lean dataset covering all miniF2F-IMO problems—including 17 official IMO problems from 2022–2023—comprising 1,329 verifiable intermediate lemmas and 45,880 lines of code. We introduce a novel fine-grained lemma decomposition technique that generates nontrivial, solvable, progressive subtasks, enabling diagnostic model evaluation. Contribution/Results: This work fills a critical gap by establishing the first open benchmark for high-difficulty formal mathematical proof, specifically designed for structured inductive reasoning. Zero-shot and few-shot evaluations on state-of-the-art models—including LLaMA-3 and Claude-3—reveal substantially lower success rates on novel lemmas compared to human performance, exposing fundamental limitations in deep mathematical reasoning. All data, code, and evaluation tools are publicly released.

Technology Category

Application Category

📝 Abstract
Using AI to write formal proofs for mathematical problems is a challenging task that has seen some advancements in recent years. Automated systems such as Lean can verify the correctness of proofs written in formal language, yet writing the proofs in formal language can be challenging for humans and machines. The miniF2F benchmark has 20 IMO problems in its test set, yet formal proofs are available only for 6 of these problems (3 of which are only written by mathematicians). The model with best accuracy can only prove 2 of these 20 IMO problems, from 1950s and 60s, while its training set is a secret. In this work, we write complete, original formal proofs for the remaining IMO problems in Lean along with 3 extra problems from IMO 2022 and 2023. This effort expands the availability of proof currently in the public domain by creating 5,880 lines of Lean proof. The goal of the paper is to pave the way for developing AI models that can automatically write the formal proofs for all the IMO problems in miniF2F and beyond by providing an evaluation benchmark. In this pursuit, we devise a method to decompose the proofs of these problems into their building blocks, constructing a dataset of 1,329 lemmas with more than 40k lines of Lean code. These lemmas are not trivial, yet they are approachable, providing the opportunity to evaluate and diagnose the failures and successes of AI models. We evaluate the ability of the SOTA LLMs on our dataset and analyze their success and failure modes from different perspectives. Our dataset and code is available at: https://github.com/roozbeh-yz/IMO-Steps.
Problem

Research questions and friction points this paper is trying to address.

Develop AI models to write formal proofs for IMO problems.
Create a dataset of 1,329 lemmas for evaluating AI models.
Expand public domain with 5,880 lines of Lean proof code.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Developed Lean-based formal proofs for IMO problems.
Created a dataset of 1,329 lemmas for AI evaluation.
Evaluated SOTA LLMs on decomposed proof building blocks.
🔎 Similar Papers