In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores

📅 2026-04-21
🏛️ arXiv.org
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Current evaluations of fairness in large language models rely excessively on standardized question-answering benchmarks, which are susceptible to prompt engineering artifacts and fail to capture authentic interactive behaviors. This work proposes MAC-Fairness, a novel framework that embeds identity variables into multi-agent natural conversations, using standardized test items as conversational seeds rather than direct evaluation instruments. Through controlled variable design, large-scale conversation simulation (8 million dialogue logs), and cross-identity behavioral metrics, the framework quantifies disparities in stance consistency and peer acceptance. Experiments reveal stable, generalizable patterns of differential treatment by models, whereas traditional benchmarks exhibit high score variance and inconsistent model rankings due to prompt sensitivity, thereby significantly advancing beyond prevailing fairness evaluation paradigms.
📝 Abstract
LLM fairness should be evaluated through in-situ conversational behavior rather than standardized-test Q&A benchmarks. We show that the standardized-test paradigm can be structurally unreliable: surface-level prompt construction choices, although entirely orthogonal to the fairness question being tested, account for the majority of score variance, shift fairness conclusions in both the direction and the magnitude, and result in severe discordance in model rankings. We develop MAC-Fairness, a multi-agent conversational framework that embeds controlled variation factors into multi-round dialogue for in-situ behavior evaluation, examining how models'conversational behavior shifts when identity is varied as part of natural multi-agent interaction. Repurposing standardized-test questions as conversation seeds rather than as the evaluation instrument, we evaluate position persistence (how they hold positions, from the self-perspective) and peer receptiveness (how receptive they are to peers, from the other-perspective) across 8 million conversation transcripts spanning multiple models and identity presence configurations. In-situ behavioral evaluation reveals stable, model-specific behavioral signatures that could generalize across benchmarks differing in fairness targets and evaluation methodologies, a form of evidence the standardized-test paradigm does not offer.
Problem

Research questions and friction points this paper is trying to address.

LLM fairness
standardized-test benchmarks
in-situ behavioral evaluation
disparate treatment
prompt construction
Innovation

Methods, ideas, or system contributions that make the work stand out.

in-situ behavioral evaluation
LLM fairness
MAC-Fairness
disparate-treatment behavior
multi-agent conversation
🔎 Similar Papers