M-IFEval: Multilingual Instruction-Following Evaluation
Existing IFEval benchmarks are English-only, limiting assessment of large language models’ (LLMs) instruction-following capabilities in multilingual and cross-cultural contexts. To address this, we introduce ML-IFEval—the first multilingual instruction-following evaluation benchmark covering French, Japanese, and Spanish. Methodologically, we extend the deterministic, rule-driven IFEval framework to multilingual settings by proposing language-adapted instruction design, cross-lingual consistency verification, and human-AI collaborative validation. Our approach integrates multilingual template engineering with automated rule-based evaluation to ensure objectivity and reproducibility. Experiments across eight state-of-the-art LLMs reveal substantial inter-lingual performance disparities, underscoring the necessity of multilingual evaluation. ML-IFEval thus provides the first open-source, subjectivity-free, and fully reproducible benchmark for internationalized LLM assessment, enabling rigorous, culture-aware evaluation of instruction following across languages.