Robust Fake News Detection using Large Language Models under Adversarial Sentiment Attacks
This study addresses the vulnerability of existing fake news detection models to adversarial sentiment manipulation, which undermines prediction stability. The work proposes AdSent, the first systematic framework leveraging large language models (LLMs) to generate sentiment-controllable adversarial examples for fake news detection. AdSent introduces a sentiment-agnostic training strategy to build detectors robust to sentiment shifts, effectively mitigating the inherent bias toward neutral sentiment prevalent in current models. Evaluated on three benchmark datasets, the proposed method significantly outperforms state-of-the-art approaches, demonstrating marked improvements in prediction consistency and accuracy on both original and sentiment-manipulated news, as well as enhanced cross-dataset generalization capability.