🤖 AI Summary
This study investigates whether Chinese-developed vision-language models exhibit alignment with national stances on politically sensitive content and examines their evolving tendency from explicit refusal to implicit reconstruction. To this end, the authors construct a balanced evaluation benchmark comprising 200 core items and seven visually abstract variants, testing nine models under multilingual and multi-prompt paradigms. They introduce a novel decomposition of multimodal censorship into two distinct signals—“refusal” and “reconstruction”—and employ a fine-grained six-dimensional auditing framework combining large language models and human expert judgments. Findings reveal that Chinese-language prompts increase reconstruction likelihood by approximately threefold; Chinese models demonstrate significantly higher reconstruction rates than non-Chinese counterparts; reconstruction is strongest in purely textual political commentary; and the Qwen model series shows a clear trend across versions of decreasing refusal and increasing reconstruction.
📝 Abstract
State-aligned distortion has been documented in China-origin text-based large language models (LLMs), but whether, and in what form, it arises in multimodal systems has not been systematically examined. We construct a balanced benchmark of 200 core entries spanning ten politically sensitive topics, plus a seven-variant visual-abstraction probe, and run nine vision-language models (VLMs), seven China-origin and two non-China, across four elicitation paradigms and two prompt languages, yielding 21,708 trials. Each response is audited on six dimensions -- explicit refusal, information integrity, visual grounding, state-aligned framing, language consistency, and response length -- by two independent frontier LLM judges, validated against three human experts on a 200-trial sample. Measuring each dimension separately lets us decompose multimodal censorship into individual signals rather than a single refusal-based score; in particular, refusal and framing are measured independently, so a model can stop refusing while still reframing. We find that (i) Chinese-language prompting roughly triples the odds of state-aligned framing, within every model; (ii) China-origin models reframe more than non-China models (direction robust across judges and human raters; magnitude 1.6--3.2x); (iii) the effect is strongest in text-only political commentary (36.5%) and is gated by recognition of the depicted subject rather than pixel detail, persisting even at silhouette for iconic images; and (iv) across four Qwen generations, state-aligned framing rises while explicit refusal falls: censorship migrates from a visible act (refusal) to an invisible one (fluent reframing). We argue this shift to invisible reframing is fundamentally a problem of human-AI interaction: it removes the very signal users rely on to recognize that information has been withheld.