Templated or fully Synthetic? Prompt construction as a confound in measuring LLM political stance beyond writing assistance

📅 2026-08-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Current approaches to assessing political stance often rely on templated prompts, which struggle to capture the nuanced dynamics of authentic dialogue and are susceptible to human manipulation, thereby introducing measurement bias. To enhance ecological validity, this work extends the IssueBench framework by introducing fully synthetic prompts generated by large language models (LLMs) from real-world prompt seeds, tailored for both information-seeking and opinion-sharing tasks. The study presents the first systematic comparison among real, templated, and fully synthetic prompts in terms of ecological validity and stance clarity. Findings reveal that templated prompts systematically overestimate model stance extremity in neutral contexts. Both human and LLM evaluations demonstrate that fully synthetic prompts achieve authenticity comparable to real prompts and significantly outperform templated ones.
📝 Abstract
Political stance detection in LLMs has long been dominated by closed-ended, multiple-choice political survey questions---originally designed for humans, and thus lacks the realism and nuance of human-AI interactions in the wild, while also being susceptible to sandbagging. The recent IssueBench framework substantially mitigates these limitations with templated prompts anchored in real-world chat logs. Given the rise in non-work-related use of GenAI assistants, we extend IssueBench beyond writing assistance to include two additional tasks, information seeking and opinion sharing. We argue that templated prompts still lack the nuance of real ones, especially for open-ended tasks, and remain recognisable as evaluation artefacts. We propose the use of fully synthetic (LLM-generated) prompts, produced under detailed instructions with real prompts as seeds. We assess the ecological validity of real, templated, and LLM-generated prompts in a small-scale study covering 3 highly contested policy issues and 3 recent geopolitical conflicts. Human and LLM annotators rank LLM-generated prompts as no less realistic than real ones and clearly more realistic than templated ones, and find that they carry their intended intent and stance more clearly; the LLMs separate templated prompts from the other two far more sharply than the humans do. In a case study, templated and LLM-generated prompts yield systematically different stance estimates for the same model, most visibly under neutral framings, where templated prompts overstate the model's leaning in the direction encoded by the topic-and-stance text (filler) slotted into their templates.
Problem

Research questions and friction points this paper is trying to address.

political stance detection
prompt construction
ecological validity
large language models
synthetic prompts
Innovation

Methods, ideas, or system contributions that make the work stand out.

synthetic prompts
political stance detection
ecological validity
LLM evaluation
IssueBench