Navigating the digital spectrum: Assessing political bias, stability, and downstream fairness in Large Language Models

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过引入稳健的政治罗盘测试框架,评估大型语言模型的政治倾向、稳定性和下游公平性问题,揭示了指令表述、语言和回答格式等因素对结果的影响。
📝 Abstract
Large Language Models are increasingly deployed as information intermediaries, yet measuring their political behavior remains fragile because questionnaire results mix model dispositions with measurement artifacts and response-elicitation biases. We introduce a robust Political Compass Test evaluation framework that samples 300 configurations across an eight-dimensional perturbation space varying language, framing, instructions, answer format, option order, and persona wording. We evaluate eight Gemma 3 and Qwen 3 models across 14 languages and three quantization levels, obtaining design-averaged political coordinates with quantified uncertainty. Most models lean Libertarian-Left on average, but instruction phrasing, language, and answer format significantly affect recovered coordinates. Cross-lingual differences primarily reflect coordinate drift rather than distinct cultural reasoning. Reverse-engineering the test also exposes axis-weighting imbalances and the collapse of degenerate responses toward the center, so near-origin estimates for the smallest models can reflect weak signal rather than centrism. Free-text reasoning and chat-then-classify elicitation alter recovered coordinates, and larger models show clearer persona separation, with a specific failure of the Authoritarian-Left persona to move most models in the intended social direction. In downstream tasks, persona effects are modest relative to model size and target group for hate-speech detection, while base and centrist prompts give the highest agreement for topic-level sentiment. Political role prompting therefore has measurable but task- and dataset-specific downstream effects.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Political Bias
Stability
Downstream Fairness
Information Intermediaries
Innovation

Methods, ideas, or system contributions that make the work stand out.

Political Compass Test
perturbation space
cross-lingual differences
downstream fairness
persona effects
🔎 Similar Papers
No similar papers found.