MedAraBench: Large-Scale Arabic Medical Question Answering Dataset and Benchmark
This work addresses the longstanding scarcity of high-quality open-source datasets and evaluation benchmarks for Arabic in medical natural language processing, which has hindered the development of multilingual large language models. To bridge this gap, the authors introduce MedAraBench, the first large-scale Arabic medical question-answering benchmark spanning 19 medical specialties and 5 difficulty levels. Data quality is ensured through manual digitization of professional medical materials, expert annotations, and a dual validation mechanism combining “LLM-as-a-Judge” with human review. The dataset and evaluation scripts are publicly released, and comprehensive evaluations across eight leading models—including GPT-5, Gemini 2.0 Flash, and Claude 4-Sonnet—reveal critical performance bottlenecks in Arabic medical tasks, thereby filling a crucial void in healthcare AI evaluation for the Arabic language.