Analysis of LLM Bias (Chinese Propaganda&Anti-US Sentiment) in DeepSeek-R1 vs. ChatGPT o3-mini-high
This study systematically evaluates the ideological neutrality of large language models (LLMs) within U.S.–China political contexts, comparing China-aligned DeepSeek-R1 and non-China-aligned ChatGPT o3-mini-high on state propaganda and anti-American sentiment. We propose the first cross-lingual (Simplified Chinese, Traditional Chinese, English), decontextualized bias evaluation framework, constructing a 1,200-item multilingual reasoning benchmark. Evaluation combines rubric-guided GPT-4o automated scoring with double-blind human annotation. Results reveal a pronounced “invisible amplifier” effect in DeepSeek-R1: its pro-state and anti-American biases are strongest in Simplified Chinese, attenuate sharply across linguistic shifts (→ Traditional Chinese → English), and generalize beyond politics into cultural domains; ChatGPT o3-mini-high remains largely ideologically neutral. The findings expose a deep coupling between linguistic representation and geopolitical alignment, offering a novel paradigm for assessing value alignment in LLMs.