🤖 AI Summary
This study investigates whether mainstream large language models (LLMs) exhibit political discourse homogenization—i.e., uncritically reproducing Western-centric terminology while neglecting geopolitical heterogeneity—when prompted with role-based identities (e.g., “Russian expert,” “U.S. expert”) on political topics concerning China, Iran, Russia, and the United States.
Method: Leveraging role-based prompt engineering and cross-model semantic similarity analysis, we systematically evaluate outputs from GPT, Gemini, Claude, and DeepSeek.
Contribution/Results: All models demonstrate significant discourse homogenization, especially on China- and Iran-related topics, yielding highly stable yet chronologically outdated narratives. Role prompting fails to elicit viewpoint diversity. This work provides the first empirical evidence of structural perspective narrowing in pretrained LLMs’ international political storytelling, revealing a systemic bias rooted in training data and architecture. It establishes a methodological benchmark for AI bias auditing and informs context-sensitive prompt design in multilingual, geopolitically nuanced applications.
📝 Abstract
Ask your chatbot to impersonate an expert from Russia and an expert from US and query it on Chinese politics. How might the outputs differ? Or, to prepare ourselves for the worse, how might they converge? Scholars have raised concerns LLM based applications can homogenize cultures and flatten perspectives. But exactly how much does LLM generated outputs converge despite explicit different role assignment? This study provides empirical evidence to the above question. The critique centres on pretrained models regurgitating ossified political jargons used in the Western world when speaking about China, Iran, Russian, and US politics, despite changes in these countries happening daily or hourly. The experiments combine role-prompting and similarity metrics. The results show that AI generated discourses from four models about Iran and China are the most homogeneous and unchanging across all four models, including OpenAI GPT, Google Gemini, Anthropic Claude, and DeepSeek, despite the prompted perspective change and the actual changes in real life. This study does not engage with history, politics, or literature as traditional disciplinary approaches would; instead, it takes cues from international and area studies and offers insight on the future trajectory of shifting political discourse in a digital space increasingly cannibalised by AI.