GPE: Evaluating Robust Evidence Aggregation for Fact Verification under Controllable GEO-Style Poisoning

📅 2026-07-22
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the lack of controllable evidence environments in existing fact verification benchmarks, which hinders robust evaluation of models under search engine optimization (SEO)-style evidence poisoning attacks. The authors propose the first multi-domain fact verification benchmark and evaluation framework that enables precise control over evidence sources and poisoning ratios, allowing systematic assessment of mainstream verification methods under adversarial retrieval conditions. Experimental results demonstrate that the proposed framework effectively uncovers robustness degradation and efficiency trade-offs that remain undetected when evaluating solely on clean data. These findings underscore the necessity of adversarial evaluation for fact verification systems and establish a new benchmark to guide the development of more robust verification approaches.
📝 Abstract
Large language models increasingly use search tools to retrieve up-to-date information, introducing a new attack surface in which retrieved documents can be manipulated. This risk is amplified by the development of generative engine optimization, which can make selected content more likely to be retrieved, cited, and adopted by models. Existing fact-verification benchmarks and evaluation frameworks do not provide the controlled evidence environments needed to assess robustness against GEO poisoning. We therefore propose GPE, which consists of a multi-domain fact-verification benchmark and an evaluation framework for controlling evidence sources and poisoning ratios. Experiments across multiple verification methods and poisoning attacks demonstrate that GPE exposes robustness degradation and efficiency trade-offs that cannot be observed through clean evaluation alone, confirming the need to evaluate fact verification under adversarial evidence environments.
Problem

Research questions and friction points this paper is trying to address.

fact verification
GEO poisoning
evidence aggregation
robustness evaluation
adversarial evidence
Innovation

Methods, ideas, or system contributions that make the work stand out.

GEO poisoning
fact verification
evidence aggregation
robustness evaluation
controlled benchmark
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zhaoqi Wang
School of Cyberspace Science and Technology, Beijing Institute of Technology
Zijian Zhang
Zijian Zhang
Beijing Institute of Technology
AI SecurityBlockchain SystemsTor NetworksData Privacy
X
Xiaomei Yuan
School of Cyberspace Science and Technology, Beijing Institute of Technology
P
Pengtao Kou
School of Cyberspace Science and Technology, Beijing Institute of Technology
Jiamou Liu
Jiamou Liu
The University of Auckland
Social NetworksArtificial IntelligenceMachine Learning
Zhen Li
Zhen Li
Beijing Institute of Technology
Vision-and-LanguageVideo Generation
L
Liehuang Zhu
School of Cyberspace Science and Technology, Beijing Institute of Technology