FramingQA: Does the Question Shape the Answer? Measuring the Compositional Framing Effect

📅 2026-09-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过FramingQA基准测试,评估大型语言模型对问题表述的敏感度及其在不同领域内提供一致答案的能力。
📝 Abstract
We introduce FramingQA, a benchmark that measures the model sensitivity to question framing across law, medicine, finance, and robotic simulations. Large language models (LLMs) often change their responses to subtle rephrasings that align with an implied stance by users. This can leave users with advice tainted by how they happened to phrase a question rather than by the underlying facts, and the consequences are highly costly in high-stakes domains. Because in the realistic scenarios, both expert practitioners and non-expert users frequently ask LLMs questions containing incomplete or misleading assumptions, models are highly susceptible to those framings. To test this, we inject the framing bias across three nested levels: a framing-biased question phrasing (root), an injected framing-biased premise prepended to a neutral question (propositional), and a premise paired with a framing-biased question (global). Evaluating nine open models (3.8B-70B) across four families, we find that strong per-variant accuracy does not guarantee the robustness across differently phrased questions under the fixed factual information.
Problem

Research questions and friction points this paper is trying to address.

question framing
large language models
sensitivity
high-stakes domains
factual information
Innovation

Methods, ideas, or system contributions that make the work stand out.

FramingQA
Question Framing
Model Sensitivity
Large Language Models
Benchmark
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Hazel H. Kim
Department of Computer Science, University of Oxford
A
Andrew M. Bean
Oxford Internet Institute, University of Oxford; Thomson Reuters
G
Guilherme Affonso Ferreira de Camargo
Faculty of Law, University of Oxford
S
Shanyu Chauhan
Department of Mechanical Engineering, University of California San Diego
Felix Drinkall
Felix Drinkall
University of Oxford
J
Jade Kosché
Faculty of Law, University of Oxford
Chenyang Ma
Chenyang Ma
University of Oxford
Embodied AIVLAAgents
G
Glory Nwaugbala
Faculty of Law, University of Oxford
Nabeel Seedat
Nabeel Seedat
University of Cambridge
Machine LearningUncertainty QuantificationData-Centric AILarge Language ModelsAI for health
B
Bradley Max Segal
Institute of Biomedical Engineering, University of Oxford
S
Samuel Recht
Department of Experimental Psychology, University of Oxford
Hinrich Schütze
Hinrich Schütze
University of Munich
natural language processing
P
Philip H. S. Torr
Department of Engineering Science, University of Oxford