EuraGovExam: A Multilingual Multimodal Benchmark from Real-World Civil Service Exams

📅 2026-03-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the absence of evaluation benchmarks for vision-language models in real-world multilingual, multimodal civil service examination scenarios. The authors introduce a novel benchmark derived from authentic civil service exams across five Eurasian regions, comprising over 8,000 high-resolution scanned questions. For the first time, the benchmark uses full-page exam document images as input, integrating multilingual text, tables, and complex layouts while emphasizing cultural authenticity and visual complexity. High-fidelity OCR preserves original document structure, and a standardized instruction format enables joint vision–language modeling and layout-aware reasoning. Experimental results reveal that even state-of-the-art models achieve only 86% accuracy, underscoring the benchmark’s difficulty and its value for advancing e-governance, public document analysis, and equitable test preparation.

Technology Category

Application Category

📝 Abstract
We present EuraGovExam, a multilingual and multimodal benchmark sourced from real-world civil service examinations across five representative Eurasian regions: South Korea, Japan, Taiwan, India, and the European Union. Designed to reflect the authentic complexity of public-sector assessments, the dataset contains over 8,000 high-resolution scanned multiple-choice questions covering 17 diverse academic and administrative domains. Unlike existing benchmarks, EuraGovExam embeds all question content--including problem statements, answer choices, and visual elements--within a single image, providing only a minimal standardized instruction for answer formatting. This design demands that models perform layout-aware, cross-lingual reasoning directly from visual input. All items are drawn from real exam documents, preserving rich visual structures such as tables, multilingual typography, and form-like layouts. Evaluation results show that even state-of-the-art vision-language models (VLMs) achieve only 86% accuracy, underscoring the benchmark's difficulty and its power to diagnose the limitations of current models. By emphasizing cultural realism, visual complexity, and linguistic diversity, EuraGovExam establishes a new standard for evaluating VLMs in high-stakes, multilingual, image-grounded settings. It also supports practical applications in e-governance, public-sector document analysis, and equitable exam preparation.
Problem

Research questions and friction points this paper is trying to address.

multilingual
multimodal
civil service exams
vision-language models
visual complexity
Innovation

Methods, ideas, or system contributions that make the work stand out.

multilingual multimodal benchmark
layout-aware reasoning
vision-language models
civil service exams
image-grounded reasoning
🔎 Similar Papers
No similar papers found.
J
JaeSeong Kim
Semyung University, Jecheon-si, Republic of Korea
C
Chaehwan Lim
Semyung University, Jecheon-si, Republic of Korea
S
Sang Hyun Gil
Semyung University, Jecheon-si, Republic of Korea
S
Suan Lee
Semyung University, Jecheon-si, Republic of Korea