🤖 AI Summary
This study addresses the vulnerability of LLM agents to reasoning forgery attacks and the limitations of existing content-based defenses. We propose Proof-of-Execution Memory (PoEM), a novel cryptographic verification mechanism that replaces content detection with HMAC-chained ledgers and trusted-layer write controls. This approach fundamentally prevents memory tampering and ensures immunity to text-paraphrasing attacks. Experiments across three models and scenarios demonstrate that PoEM reduces attack success rates to 0% with zero false positives, significantly outperforming SENTINEL. With microsecond-level verification overhead, PoEM effectively resolves the capability paradox wherein more powerful models exhibit greater susceptibility to attacks, thereby providing mechanism-level security guarantees for autonomous agents.
📝 Abstract
LLM agents are stateless and rely on external memory to carry context between steps. Because agents treat that memory as trustworthy, an adversary who can write to it can steer their behavior. The FARMA attack does this with no malicious command: it inserts fabricated entries into the agent's reasoning memory claiming a required safety step is already done, so the agent skips it. SENTINEL, the defense proposed with FARMA, scores entries against a fixed list of suspicious wordings; its authors note that an attacker who knows the list can reword the forgery and evade it, and leave this open. We show the gap is worse than stated. An automated attacker that simply asks a language model to reword the forgery evades SENTINEL on its first try, reducing its protection to zero on every model tested. We also find a capability paradox: the attack succeeds far more often on stronger models (98-100% on GPT-4o and GPT-4o-mini) than on Llama-3.1-8B (44%), because more capable agents follow reworded claims more faithfully, so the threat grows with capability. We propose Proof-of-Execution Memory (PoEM), which does not inspect memory at all. PoEM keeps a separate, tamper-evident, HMAC-chained ledger of the safety steps that actually executed, writable only by the trusted action layer, and allows a skip only if the ledger confirms real execution. An attacker can change what memory says but cannot forge a ledger entry for a step that never ran, so rewording no longer helps. Across three models and three scenarios, PoEM drives attack success to 0% while leaving legitimate operation intact (0% false positives in eight of nine cells, 1.7% in the ninth, within sampling noise), whereas SENTINEL wrongly blocks 33-50% of legitimate operations. PoEM also withstands attacks aimed at itself, adds microseconds of overhead, and works unchanged in a real LangChain agent. PoEM protects exactly the decisions it gates.