🤖 AI Summary
This work addresses the limited cultural evaluation of large language models, which has predominantly focused on high-resource standard languages while overlooking regional cultures and dialectal communities. To bridge this gap, the authors introduce BavGround, the first multilingual benchmark tailored to Bavarian regional culture, encompassing English, German, and Bavarian. The benchmark comprises 206 multiple-choice questions across eight cultural domains, yielding 618 parallel multilingual samples. The study employs diverse evaluation protocols—including answer-letter scoring, option-text likelihood, and semantic matching—and includes continual pretraining analyses. Evaluations across 16 prominent models reveal that strong multilingual models achieve the best overall performance, yet still exhibit significant deficiencies in Bavarian language comprehension and culturally grounded reasoning. Notably, model rankings vary substantially depending on the chosen evaluation protocol.
📝 Abstract
Cultural evaluation of large language models (LLMs) often focuses on high-resource standard languages, leaving regional culture and dialect communities underrepresented. We introduce BavGround, a benchmark for evaluating Bavarian regional cultural grounding and dialect competence across English, German and Bavarian. BavGround contains 206 multiple-choice source questions across eight cultural domains per language, yielding 618 multi-parallel instances, with items covering both broadly accessible cultural knowledge and source-grounded regional knowledge from journalism, historical sources, and specialist literature. We evaluate fifteen 7B-10B open-weight instruction-tuned models and one closed-model reference. Strong multilingual models perform best overall, but performance drops on Bavarian items and source-grounded questions, indicating persistent difficulty with dialectal and localized cultural knowledge. We further show that conclusions depend strongly on evaluation protocol: raw answer-letter scoring, shuffled-letter scoring, option-text likelihood, generated-answer parsing, and semantic matching can produce different absolute scores and rankings, especially for regionally adapted models. Finally, an exploratory analysis of GENBA-10B checkpoints suggests that continued pretraining improves answer-content likelihood unevenly across domains, while dialect competence remains comparatively weak. BavGround supports localized, protocol-aware evaluation of cultural representation in LLMs.