MEMORY Wins All: Indirect Bias Injection Attacks via Social Media Feeds

📅 2026-08-22
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究提出了一种通过社交媒体内容间接注入偏见的攻击方法IBIA,影响AI代理的行为,并评估了其有效性及防御措施。
📝 Abstract
Personal AI agents routinely consume external content while performing tasks such as web browsing, email processing, and SNS feed summarization, and they retain selected information or execution results in persistent memory for later use. We show that this ordinary ingestion of external content opens an indirect path for manipulating subsequent agent behavior. Based on this observation, we present IBIA, an Indirect Bias Injection Attack that plants an adversary-aligned stance on a specific topic into a victim agent's memory through external content, without direct access to the agent, its memory, or future user queries. For this, IBIA combines three mechanisms: comment cloaking, which keeps the crafted content consistent with the surrounding discussion, comment watermarking, which enables lightweight identification during curation, and category anchoring, which makes the retained stance salient under later related requests. We evaluate IBIA on BiasBench, a benchmark of 6,000 adversary-crafted social comments and 120 email instances. The watermark-based curation identifies 95.9% of the injected comments. Under the OpenClaw setting, IBIA achieves adversary-aligned response rates (AARs) of 91.2% on average across four downstream tasks, including 86.6% on the frontier GPT-5.5. We further propose a memory boundary defense that detects the injected bias and reduces AARs to 80.6%.
Problem

Research questions and friction points this paper is trying to address.

Indirect Bias Injection
Social Media Feeds
Personal AI Agents
External Content Ingestion
Behavior Manipulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Indirect Bias Injection Attack
Comment Cloaking
Category Anchoring
🔎 Similar Papers
No similar papers found.