Analysis of LLM Bias (Chinese Propaganda&Anti-US Sentiment) in DeepSeek-R1 vs. ChatGPT o3-mini-high

📅 2025-06-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study systematically evaluates the ideological neutrality of large language models (LLMs) within U.S.–China political contexts, comparing China-aligned DeepSeek-R1 and non-China-aligned ChatGPT o3-mini-high on state propaganda and anti-American sentiment. We propose the first cross-lingual (Simplified Chinese, Traditional Chinese, English), decontextualized bias evaluation framework, constructing a 1,200-item multilingual reasoning benchmark. Evaluation combines rubric-guided GPT-4o automated scoring with double-blind human annotation. Results reveal a pronounced “invisible amplifier” effect in DeepSeek-R1: its pro-state and anti-American biases are strongest in Simplified Chinese, attenuate sharply across linguistic shifts (→ Traditional Chinese → English), and generalize beyond politics into cultural domains; ChatGPT o3-mini-high remains largely ideologically neutral. The findings expose a deep coupling between linguistic representation and geopolitical alignment, offering a novel paradigm for assessing value alignment in LLMs.

Technology Category

Application Category

📝 Abstract
Large language models (LLMs) increasingly shape public understanding and civic decisions, yet their ideological neutrality is a growing concern. While existing research has explored various forms of LLM bias, a direct, cross-lingual comparison of models with differing geopolitical alignments-specifically a PRC-system model versus a non-PRC counterpart-has been lacking. This study addresses this gap by systematically evaluating DeepSeek-R1 (PRC-aligned) against ChatGPT o3-mini-high (non-PRC) for Chinese-state propaganda and anti-U.S. sentiment. We developed a novel corpus of 1,200 de-contextualized, reasoning-oriented questions derived from Chinese-language news, presented in Simplified Chinese, Traditional Chinese, and English. Answers from both models (7,200 total) were assessed using a hybrid evaluation pipeline combining rubric-guided GPT-4o scoring with human annotation. Our findings reveal significant model-level and language-dependent biases. DeepSeek-R1 consistently exhibited substantially higher proportions of both propaganda and anti-U.S. bias compared to ChatGPT o3-mini-high, which remained largely free of anti-U.S. sentiment and showed lower propaganda levels. For DeepSeek-R1, Simplified Chinese queries elicited the highest bias rates; these diminished in Traditional Chinese and were nearly absent in English. Notably, DeepSeek-R1 occasionally responded in Simplified Chinese to Traditional Chinese queries and amplified existing PRC-aligned terms in its Chinese answers, demonstrating an"invisible loudspeaker"effect. Furthermore, such biases were not confined to overtly political topics but also permeated cultural and lifestyle content, particularly in DeepSeek-R1.
Problem

Research questions and friction points this paper is trying to address.

Compare ideological bias in PRC-aligned vs non-PRC LLMs
Measure propaganda and anti-US sentiment in Chinese-language outputs
Analyze how translation affects bias expression across languages
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-lingual comparison of PRC vs. non-PRC LLM biases
Hybrid evaluation pipeline with GPT-4o and human annotation
De-contextualized corpus for measuring propaganda and sentiment
💼 Related Jobs
No related jobs found.
P
PeiHsuan Huang
Taiwan AI Labs, Taipei, Taiwan
Z
ZihWei Lin
Taiwan AI Labs, Taipei, Taiwan
S
Simon Imbot
Taiwan AI Labs, Taipei, Taiwan
W
WenCheng Fu
National Defense University, Taipei, Taiwan
Ethan Tu
Ethan Tu
Taiwan AI Labs, Taipei, Taiwan