AI4Math: A Native Spanish Benchmark for University-Level Mathematical Reasoning in Large Language Models

📅 2025-05-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing mathematical reasoning benchmarks are predominantly English-based or translation-derived, risking semantic drift and obscuring language-specific reasoning deficiencies. Method: We introduce AI4Math—the first natively authored Spanish-language university-level mathematical reasoning benchmark—comprising 105 original problems spanning seven domains (e.g., algebra, calculus, geometry) with expert-annotated step-by-step solutions. It enables robust cross-lingual analysis without translation artifacts. We conduct bilingual (Spanish/English) zero-shot and chain-of-thought evaluations on state-of-the-art models including GPT-4o, DeepSeek-R1/V3, DeepSeek-o3 mini, and LLaMA-3.3. Contribution/Results: Top-performing models achieve >70% accuracy on Spanish tasks; notably, GPT-4o attains higher zero-shot accuracy in Spanish than in English. However, performance drops sharply (<40%) on geometry, combinatorics, and probability problems—revealing persistent domain-specific weaknesses. AI4Math thus provides a linguistically grounded, domain-diverse resource for probing multilingual mathematical reasoning capabilities.

Technology Category

Application Category

📝 Abstract
Existing mathematical reasoning benchmarks are predominantly English only or translation-based, which can introduce semantic drift and mask languagespecific reasoning errors. To address this, we present AI4Math, a benchmark of 105 original university level math problems natively authored in Spanish. The dataset spans seven advanced domains (Algebra, Calculus, Geometry, Probability, Number Theory, Combinatorics, and Logic), and each problem is accompanied by a step by step human solution. We evaluate six large language models GPT 4o, GPT 4o mini, o3 mini, LLaMA 3.3 70B, DeepSeek R1 685B, and DeepSeek V3 685B under four configurations: zero shot and chain of thought, each in Spanish and English. The top models (o3 mini, DeepSeek R1 685B, DeepSeek V3 685B) achieve over 70% accuracy, whereas LLaMA 3.3 70B and GPT-4o mini remain below 40%. Most models show no significant performance drop between languages, with GPT 4o even performing better on Spanish problems in the zero shot setting. Geometry, Combinatorics, and Probability questions remain persistently challenging for all models. These results highlight the need for native-language benchmarks and domain-specific evaluations to reveal reasoning failures not captured by standard metrics.
Problem

Research questions and friction points this paper is trying to address.

Native Spanish benchmark for university-level math reasoning in LLMs
Evaluates LLMs on original Spanish math problems across seven domains
Reveals language-specific reasoning errors and domain challenges in LLMs
Innovation

Methods, ideas, or system contributions that make the work stand out.

Native Spanish math benchmark AI4Math
Evaluates six LLMs in Spanish and English
Focuses on university-level math domains
🔎 Similar Papers
No similar papers found.
M
Miguel Angel Penaloza Perez
Centro de Investigacion Cientifica y de Educacion Superi Mexico
B
Bruno Lopez Orozco
Facultad de Ciencias Unam Mexico
J
Jesus Tadeo Cruz Soto
Facultad de Matematicas Universidad Veracruzana Mexico
M
Michelle Bruno Hernandez
Carreras con Impacto
M
Miguel Angel Alvarado Gonzalez
Carreras con Impacto
S
Sandra Malagon Carreras con Impacto
Carreras con Impacto