Investigating Bias: A Multilingual Pipeline for Generating, Solving, and Evaluating Math Problems with LLMs

📅 2025-09-22
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study identifies significant language bias in large language models (LLMs) for educational support: the same mathematical problem yields systematically lower solution quality in Arabic and German compared to English, with Arabic exhibiting the weakest performance. Method: We introduce the first multilingual framework for K–10 mathematics in Germany—supporting automated problem generation, step-by-step solving, and LLM-driven cross-lingual evaluation—integrating GPT-4o-mini, Gemini 2.5 Flash, Qwen-plus, and Claude 3.5 Haiku. Using 628 original problems, we conduct end-to-end solving and adjudicated comparative evaluation across English, German, and Arabic. Contribution/Results: Results confirm statistically significant superiority of English solutions over German and Arabic counterparts; critically, this disparity persists even when controlling for translation artifacts, indicating intrinsic imbalances in LLMs’ multilingual reasoning capabilities. This work provides the first quantitative validation of language bias in foundational mathematics education contexts and establishes a reproducible methodology for multilingual educational assessment.

Technology Category

Application Category

📝 Abstract
Large Language Models (LLMs) are increasingly used for educational support, yet their response quality varies depending on the language of interaction. This paper presents an automated multilingual pipeline for generating, solving, and evaluating math problems aligned with the German K-10 curriculum. We generated 628 math exercises and translated them into English, German, and Arabic. Three commercial LLMs (GPT-4o-mini, Gemini 2.5 Flash, and Qwen-plus) were prompted to produce step-by-step solutions in each language. A held-out panel of LLM judges, including Claude 3.5 Haiku, evaluated solution quality using a comparative framework. Results show a consistent gap, with English solutions consistently rated highest, and Arabic often ranked lower. These findings highlight persistent linguistic bias and the need for more equitable multilingual AI systems in education.
Problem

Research questions and friction points this paper is trying to address.

Evaluating linguistic bias in LLMs' math problem-solving across languages
Developing multilingual pipeline for generating and assessing math exercises
Analyzing solution quality disparities in English, German, and Arabic
Innovation

Methods, ideas, or system contributions that make the work stand out.

Automated multilingual pipeline for math problem generation
LLM-based step-by-step solution evaluation with comparative framework
Held-out panel of LLM judges assessing solution quality
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Mariam Mahran
HTW Berlin University of Applied Sciences, Treskowallee 8, 10318 Berlin, Germany
Katharina Simbeck
Katharina Simbeck
HTW Berlin