Eroding the Truth-Default: A Causal Analysis of Human Susceptibility to Foundation Model Hallucinations and Disinformation in the Wild

📅 2026-01-30
📈 Citations: 2
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the growing threat posed by increasingly realistic misinformation generated by foundation models to trustworthy online information ecosystems. The authors propose a dual-axis framework, JudgeGPT and RogueGPT, which decouples “factual accuracy” from “source attribution,” and introduces the concept of the “fluency trap” to elucidate the cognitive mechanisms underlying human susceptibility to hallucinations. Leveraging a structural causal model, 918 human evaluations, and comparisons across multiple models—including GPT-4 and Llama-2—the research finds that political orientation exerts minimal influence, whereas familiarity with fake news serves as a key mediating variable (r = 0.35). Notably, GPT-4–generated content achieves a human–machine confusion rate of 0.20. The findings advocate for “prebunking” interventions centered on cognitive source monitoring, offering both empirical grounding and a novel pathway toward fostering a more reliable information ecosystem.

Technology Category

Application Category

📝 Abstract
As foundation models (FMs) approach human-level fluency, distinguishing synthetic from organic content has become a key challenge for Trustworthy Web Intelligence. This paper presents JudgeGPT and RogueGPT, a dual-axis framework that decouples"authenticity"from"attribution"to investigate the mechanisms of human susceptibility. Analyzing 918 evaluations across five FMs (including GPT-4 and Llama-2), we employ Structural Causal Models (SCMs) as a principal framework for formulating testable causal hypotheses about detection accuracy. Contrary to partisan narratives, we find that political orientation shows a negligible association with detection performance ($r=-0.10$). Instead,"fake news familiarity"emerges as a candidate mediator ($r=0.35$), suggesting that exposure may function as adversarial training for human discriminators. We identify a"fluency trap"where GPT-4 outputs (HumanMachineScore: 0.20) bypass Source Monitoring mechanisms, rendering them indistinguishable from human text. These findings suggest that"pre-bunking"interventions should target cognitive source monitoring rather than demographic segmentation to ensure trustworthy information ecosystems.
Problem

Research questions and friction points this paper is trying to address.

foundation models
hallucinations
disinformation
human susceptibility
trustworthy AI
Innovation

Methods, ideas, or system contributions that make the work stand out.

Structural Causal Models
fluency trap
source monitoring
fake news familiarity
foundation model hallucinations
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Alexander Loth
Frankfurt University of Applied Sciences, Frankfurt am Main, Germany
M
Martin Kappes
Frankfurt University of Applied Sciences, Frankfurt am Main, Germany
M
Marc-Oliver Pahl
IMT Atlantique, UMR IRISA, Chaire Cyber CNI, Rennes, France