MathArena: Evaluating LLMs on Uncontaminated Math Competitions
Public benchmark datasets (e.g., AIME 2024) suffer from widespread data leakage, confounding LLM mathematical reasoning evaluation with memorization effects. Method: We introduce the first contamination-free benchmark for mathematical reasoning, built on real-time released contest problems—149 unseen questions from five major competitions (AIME, SMT, USAMO, etc.)—governed by a strict decontamination protocol aligned with official contest release windows. Contribution/Results: We propose a novel multi-granularity scoring scheme jointly evaluating answer correctness and proof rigor, enabling the first standardized, systematic assessment of formal proof generation. Experiments reveal that while state-of-the-art models excel on uncontaminated problem-solving tasks (e.g., SMT 2025), their scores drop below 25% on USAMO 2025 proof-generation tasks—unambiguously exposing deficiencies in formal deductive reasoning. This stark performance gap validates the benchmark’s efficacy in disentangling genuine reasoning capability from dataset memorization.