From Global Benchmarks to Local Evaluations: Benchmarking LLMs for the German Public Sector

📅 2026-08-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对德语公共部门选择合适的大语言模型问题,提出MÖVE评估框架,从能耗、透明度和政党立场知识三个维度进行综合评价。
📝 Abstract
Public institutions face a persistent challenge in selecting LLMs suited to their specific context. Existing benchmarks, however, are of limited use as they primarily reflect English-language and US-centric settings, and often only evaluate task performance. In this paper, we present first results of MÖVE, a holistic evaluation framework for the German public sector, examining three rarely considered governance dimensions: energy consumption, provider transparency, and knowledge of German-party positions. Our results reveal significant trade-offs, with no single model excelling across all dimensions: estimated energy consumption varies more than 60-fold and is not explained by model size alone, information disclosure varies systematically across providers, and European models do not exhibit stronger knowledge of German party positions. Model selection for public institutions thus cannot rely on performance rankings alone. Instead, evaluations should also reflect the governance requirements of the deployment context.
Problem

Research questions and friction points this paper is trying to address.

Public Institutions
LLM Selection
Benchmark Limitations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Governance Dimensions
Energy Consumption
Provider Transparency
German-party Positions
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
C
Camilla Dalerci
Innovations Department, Bundesdruckerei GmbH, Berlin, Germany
Thilo Michael
Thilo Michael
TU Berlin
R
Robin Schaefer
Innovations Department, Bundesdruckerei GmbH, Berlin, Germany
Daniel Weinland
Daniel Weinland
Bundesdruckerei GmbH
Machine LearningNLPComputer Vision