🤖 AI Summary
This study addresses the growing threat posed by increasingly realistic misinformation generated by foundation models to trustworthy online information ecosystems. The authors propose a dual-axis framework, JudgeGPT and RogueGPT, which decouples “factual accuracy” from “source attribution,” and introduces the concept of the “fluency trap” to elucidate the cognitive mechanisms underlying human susceptibility to hallucinations. Leveraging a structural causal model, 918 human evaluations, and comparisons across multiple models—including GPT-4 and Llama-2—the research finds that political orientation exerts minimal influence, whereas familiarity with fake news serves as a key mediating variable (r = 0.35). Notably, GPT-4–generated content achieves a human–machine confusion rate of 0.20. The findings advocate for “prebunking” interventions centered on cognitive source monitoring, offering both empirical grounding and a novel pathway toward fostering a more reliable information ecosystem.
📝 Abstract
As foundation models (FMs) approach human-level fluency, distinguishing synthetic from organic content has become a key challenge for Trustworthy Web Intelligence. This paper presents JudgeGPT and RogueGPT, a dual-axis framework that decouples"authenticity"from"attribution"to investigate the mechanisms of human susceptibility. Analyzing 918 evaluations across five FMs (including GPT-4 and Llama-2), we employ Structural Causal Models (SCMs) as a principal framework for formulating testable causal hypotheses about detection accuracy. Contrary to partisan narratives, we find that political orientation shows a negligible association with detection performance ($r=-0.10$). Instead,"fake news familiarity"emerges as a candidate mediator ($r=0.35$), suggesting that exposure may function as adversarial training for human discriminators. We identify a"fluency trap"where GPT-4 outputs (HumanMachineScore: 0.20) bypass Source Monitoring mechanisms, rendering them indistinguishable from human text. These findings suggest that"pre-bunking"interventions should target cognitive source monitoring rather than demographic segmentation to ensure trustworthy information ecosystems.