How Much Can We Trust LLM Search Agents? Measuring Endorsement Vulnerability to Web Content Manipulation
This work addresses the vulnerability of large language model (LLM) search agents to adversarial manipulation, wherein attacker-controlled web content may be erroneously treated as credible evidence, leading LLMs to endorse harmful claims. To systematically evaluate this risk, the authors propose SearchGEO, a novel evaluation framework that introduces recommendation reliability as a core dimension of LLM backend safety. SearchGEO establishes a controlled and reproducible paradigm for assessing endorsement vulnerabilities through an integrated pipeline comprising web evidence manipulation, five adversarial attack patterns, multi-level output metrics, and auxiliary skill probes—such as command conversion. Evaluations across 13 mainstream LLMs on 308 cases reveal attack success rates ranging from 0.0% (Claude-Sonnet-4.6) to 31.4% (Gemini-3-Flash), with substantial response variation even among models of similar architecture; auxiliary probes further indicate that Claude tends toward excessive refusal, whereas GPT models exhibit undue trust.