From Values to Benchmarks: Evaluating Large Language Models for Governmental Use in Dutch

πŸ“… 2026-08-10
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the gap in existing large language model (LLM) evaluation frameworks, which often neglect public administration values and Dutch linguistic characteristics. Through expert consultation, user surveys, and civil servant interviews, the authors develop β€œGrip on LLMs,” a systematic evaluation framework tailored for Dutch government applications. It assesses over thirty general-purpose and Dutch-specific models across six dimensions: factuality, honesty, social bias, energy consumption, cost, and training data transparency. Innovatively integrating public governance principles with local language requirements, the work introduces a multi-dimensional trade-off perspective, revealing that factuality and honesty are governed by distinct mechanisms. A visualization tool is also developed to support non-technical decision-makers in model selection. Findings indicate no single model dominates across all criteria; high performance frequently entails greater environmental and economic costs, and bias shows no significant correlation with model capability. The results are publicly released as an accessible, multi-stakeholder model overview platform.
πŸ“ Abstract
Large language models are increasingly being deployed in governmental settings, yet few existing evaluation frameworks jointly reflect the values of public administration and the linguistic requirements of non-English contexts. We present the "Grip on LLMs" framework, a systematic evaluation suite for Dutch governmental use developed in collaboration with domain experts from a major Dutch municipal organisation. Through an advisory board process, user research, and a survey of the users of a civil-servant chatbot, we identify six evaluation dimensions (factuality, honesty, social bias, energy consumption, cost, and training data transparency) and operationalise them into a benchmark suite covering more than 30 multilingual and Dutch-specific models. Our results reveal that no single model excels across all dimensions, and that trade-offs are unavoidable: higher quality consistently comes at greater environmental impact and financial cost, while bias remains largely independent of both. We further find that factuality (whether a model answers correctly) and honesty (whether a model acknowledges what it does not know) are governed by distinct properties, with high factuality not implying high honesty. To make these findings actionable for non-technical audiences, we release a publicly accessible, user-friendly model overview designed for the full range of stakeholders involved in governmental LLM selection, from engineers to policymakers.
Problem

Research questions and friction points this paper is trying to address.

large language models
governmental use
evaluation framework
public administration values
non-English contexts
Innovation

Methods, ideas, or system contributions that make the work stand out.

evaluation framework
governmental LLMs
Dutch language
factuality vs. honesty
multilingual benchmark
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
L
Laurens Samson
City of Amsterdam, Socially Intelligent Artificial Systems Group, University of Amsterdam
I
Iva Gornishka
City of Amsterdam, Socially Intelligent Artificial Systems Group, University of Amsterdam
G
Gossa LΓ΄
City of Amsterdam, Socially Intelligent Artificial Systems Group, University of Amsterdam
Yuki M. Asano
Yuki M. Asano
Full Professor, Head of FunAI Lab, University of Technology Nuremberg
Deep LearningMultimodal LearningSelf-supervised LearningLarge Model AdaptationLLMs
S
Sennay Ghebreab
Socially Intelligent Artificial Systems Group, University of Amsterdam