From Values to Benchmarks: Evaluating Large Language Models for Governmental Use in Dutch
This study addresses the gap in existing large language model (LLM) evaluation frameworks, which often neglect public administration values and Dutch linguistic characteristics. Through expert consultation, user surveys, and civil servant interviews, the authors develop “Grip on LLMs,” a systematic evaluation framework tailored for Dutch government applications. It assesses over thirty general-purpose and Dutch-specific models across six dimensions: factuality, honesty, social bias, energy consumption, cost, and training data transparency. Innovatively integrating public governance principles with local language requirements, the work introduces a multi-dimensional trade-off perspective, revealing that factuality and honesty are governed by distinct mechanisms. A visualization tool is also developed to support non-technical decision-makers in model selection. Findings indicate no single model dominates across all criteria; high performance frequently entails greater environmental and economic costs, and bias shows no significant correlation with model capability. The results are publicly released as an accessible, multi-stakeholder model overview platform.