EuraGovExam: A Multilingual Multimodal Benchmark from Real-World Civil Service Exams
This work addresses the absence of evaluation benchmarks for vision-language models in real-world multilingual, multimodal civil service examination scenarios. The authors introduce a novel benchmark derived from authentic civil service exams across five Eurasian regions, comprising over 8,000 high-resolution scanned questions. For the first time, the benchmark uses full-page exam document images as input, integrating multilingual text, tables, and complex layouts while emphasizing cultural authenticity and visual complexity. High-fidelity OCR preserves original document structure, and a standardized instruction format enables joint vision–language modeling and layout-aware reasoning. Experimental results reveal that even state-of-the-art models achieve only 86% accuracy, underscoring the benchmark’s difficulty and its value for advancing e-governance, public document analysis, and equitable test preparation.