🤖 AI Summary
This study addresses the limitations of existing large language model (LLM) evaluation methods, which rely on unstructured benchmarks and struggle to causally disentangle sources of bias—such as baseline traits, contextual confounding, or interaction effects. To overcome this, the authors propose the first analytical evaluation framework grounded in precise Likert-scale response distributions, integrating psychometric principles with LLM mechanics. The framework employs a full factorial experimental design, token-level probability mass functions to eliminate sampling noise, and combines ordinal consensus metrics with distributional analysis of variance. Experiments across five mainstream LLMs successfully uncover systematic national-origin biases obscured by conventional aggregate metrics, demonstrating the framework’s high sensitivity and validity in assessing consumer ethnocentrism.
📝 Abstract
As Large Language Models (LLMs) are increasingly deployed as autonomous agents, accurately evaluating their latent values and biases is critical. The NLP community typically evaluates models using large, unstructured benchmarks. While effective for general capabilities, these datasets fundamentally conflate causal mechanisms: even when an aggregate bias is detected, unstructured evaluations cannot disentangle whether it stems from baseline traits, contextual confounders, or complex interactions. To address this, we introduce an analytically exact framework for the controlled behavioral evaluation of LLMs. We bridge human psychometrics with LLM mechanics by resolving gaps in design, measurement, and analysis. First, we replace unstructured prompting with fully crossed factorial experiments to systematically isolate causal main and interaction effects. Second, we eliminate Monte Carlo text sampling noise by operating directly on exact, token-level Probability Mass Functions (PMFs). Third, we derive a multivariate ordinal consensus metric and a distributional ANOVA to process these PMFs analytically. We validate our framework with a case study on consumer ethnocentrism across five LLMs, demonstrating how our approach isolates systemic country-of-origin biases that aggregate benchmarks otherwise obscure.