Effects of Answer Format Variation on Gender Bias in Large Language Models

📅 2026-08-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过改变回答格式(闭合、李克特量表和开放)评估大型语言模型中的性别偏见,发现不同格式显著影响偏见测量结果。
📝 Abstract
Gender bias or other social biases in large language models (LLMs) are frequently evaluated with question answering or survey benchmarks where the LLM needs to give a response in a predefined answer format. It is well known in survey science that the answer format has a substantial impact on answers, just as LLMs are sensitive to the prompt wording. However, to our knowledge it has not been studied yet how changes in answer format impact the measurement of gender bias in LLMs and their alignment with human response distributions. We evaluate three instruction-tuned models on the BBQ benchmark and OpinionQA survey data across closed-ended, Likert-scaled and open-ended formats, comparing bias measurement and distributional alignment under otherwise identical conditions. We find that answer format does substantially alter measured outcomes, including reversals in order rankings. These differences arise because each format elicits distinct response behaviours, such as forced-choice selection, scale-based distributions and refusal in free-text generation. Our findings highlight the importance of treating answer format as a substantive component of LLM evaluation and motivate multi-format designs for more robust model assessment.
Problem

Research questions and friction points this paper is trying to address.

Gender Bias
Answer Format
Large Language Models
Evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

answer format
gender bias
large language models
multi-format designs
K
Ksenia Merzlyakova
Institute for Natural Language Processing, University of Stuttgart
Sebastian Padó
Sebastian Padó
Professor of Computational Linguistics, Computer Science, Stuttgart University
Natural Language ProcessingSemanticsComputational Linguistics
F
Franziska Weeber
Institute for Natural Language Processing, University of Stuttgart