ArabicNumBench: Evaluating Arabic Number Reading in Large Language Models

📅 2026-02-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study systematically evaluates the performance of large language models on Arabic numeral reading tasks—including both Eastern and Western Arabic-Indic digits—across six contextual categories: plain numbers, addresses, dates, quantities, prices, and others. Using a comprehensive benchmark comprising 210 tasks and 59,010 test cases, the authors assess 71 models under four prompting strategies: zero-shot, zero-shot chain-of-thought (CoT), few-shot, and few-shot CoT. For the first time, numerical accuracy is disentangled from instruction-following capability, revealing a wide performance range (14.29%–99.05%) and demonstrating that few-shot CoT improves accuracy by 2.8× over zero-shot. Notably, only six models consistently produce structured outputs, highlighting that high accuracy does not guarantee reliable structured responses; thus, post-processing remains essential for practical deployment.

Technology Category

Application Category

📝 Abstract
We present ArabicNumBench, a comprehensive benchmark for evaluating large language models on Arabic number reading tasks across Eastern Arabic-Indic numerals (0-9 in Arabic script) and Western Arabic numerals (0-9). We evaluate 71 models from 10 providers using four prompting strategies (zero-shot, zero-shot CoT, few-shot, few-shot CoT) on 210 number reading tasks spanning six contextual categories: pure numerals, addresses, dates, quantities, and prices. Our evaluation comprises 59,010 individual test cases and tracks extraction methods to measure structured output generation. Evaluation reveals substantial performance variation, with accuracy ranging from 14.29\% to 99.05\% across models and strategies. Few-shot Chain-of-Thought prompting achieves 2.8x higher accuracy than zero-shot approaches (80.06\% vs 28.76\%). A striking finding emerges: models achieving elite accuracy (98-99\%) often produce predominantly unstructured output, with most responses lacking Arabic CoT markers. Only 6 models consistently generate structured output across all test cases, while the majority require fallback extraction methods despite high numerical accuracy. Comprehensive evaluation of 281 model-strategy combinations demonstrates that numerical accuracy and instruction-following represent distinct capabilities, establishing baselines for Arabic number comprehension and providing actionable guidance for model selection in production Arabic NLP systems.
Problem

Research questions and friction points this paper is trying to address.

Arabic number reading
large language models
structured output
numeral systems
instruction following
Innovation

Methods, ideas, or system contributions that make the work stand out.

Arabic number reading
structured output generation
Chain-of-Thought prompting
multilingual LLM evaluation
numeral systems
🔎 Similar Papers
No similar papers found.
A
Anas Alhumud
Saudi Data and Artificial Intelligence Authority, Saudi Arabia
A
Abdulaziz Alhammadi
Imam Mohammad Ibn Saud Islamic University, Saudi Arabia
M
Muhammad Badruddin Khan
Imam Mohammad Ibn Saud Islamic University, Saudi Arabia