Security of Foundation-Model-Powered Embodied Agents: Attack Surfaces, Attacks, Defenses, and Evaluation

📅 2026-08-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the propagation of security risks from foundation models to physical behaviors in embodied agents and the challenge of localizing attack entry points. We propose a systematic security analysis framework centered on trust boundaries, introducing the novel "First Compromised Trust Boundary" principle to decouple attack surfaces from underlying mechanisms. By partitioning the system into five layers and twelve attack surfaces, combined with quantitative analysis and cross-layer propagation assessment, we elucidate risk evolution patterns. The research identifies perception and action interfaces as critical attack hotspots while highlighting significant gaps in contextual memory, middleware, and multi-agent trust. These findings provide theoretical foundations for securing embodied intelligence and outline key open challenges for future defense mechanisms.
📝 Abstract
Foundation models are increasingly used for perception, reasoning, planning, and action generation in embodied agents, creating security risks that can propagate from digital inputs to physical behavior. Existing surveys often organize threats by mechanisms such as jailbreaks, prompt injection, backdoors, poisoning, or adversarial examples, but these categories do not consistently identify where an adversary first enters the embodied control loop. We present a trust-boundary-centric survey of foundation-model-powered embodied-agent security. Using a first-compromised-trust-boundary principle, we separate attack surface from attack mechanism and organize the system into five layers and twelve attack surfaces spanning the model supply chain, user instructions, context and memory, physical semantic environments, multimodal perception, world state, internal reasoning, task planning, action interfaces, middleware, multi-agent communication, and execution control. Based on 58 attack records and 61 defense records collected through August 15, 2026, we analyze representative attacks, cross-layer propagation, defense placement, and evaluation practices. Our quantitative analysis shows that attack research is concentrated on multimodal perception and action interfaces, while defenses are especially concentrated on action-level and runtime protection. Context and long-term memory, middleware and networking, world-state integrity, and multi-agent trust remain comparatively underexplored. We conclude with open challenges in state provenance, compositional defenses, long-horizon attack propagation, physical realizability, Byzantine multi-robot behavior, and unified closed-loop evaluation.
Problem

Research questions and friction points this paper is trying to address.

Embodied Agents
Foundation Models
Security
Trust Boundary
Attack Surface
Innovation

Methods, ideas, or system contributions that make the work stand out.

Trust Boundary
Embodied Agent Security
Attack Surface Taxonomy
Cross-layer Propagation
Closed-loop Evaluation