Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents

📅 2026-08-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对自主大语言模型代理的安全问题,提出了一种名为LoopHarness的方法,通过在循环级别恢复持久、非衰减的安全状态来解决安全状态重置导致的问题。
📝 Abstract
Large language model agents are increasingly deployed as autonomous loops. Starting from one human goal, such a system repeatedly discovers work, plans, executes tool calls, verifies outcomes and persists state across many unattended iterations. The agent safeguards in wide use, however, are defined over a single trajectory, and their safety state is re-initialized when the next trajectory begins. We show that this is a failure of composition rather than an implementation detail. Our central result is a separation: against an attack whose evidence is fragmented across several iterations, every trajectory-scoped monitor has a true-positive rate equal to its false-positive rate, however expressive it is, because the evidence it would need never appears in the window it sees, whereas a monitor retaining cross-iteration state separates the two perfectly. We further show that the obvious repair of carrying a geometrically decaying risk score is insufficient, because the cooling-off period a patient adversary must wait is a constant that does not grow with the horizon $N$. We then present LoopHarness, which restores a persistent, non-decaying safety state at the loop level. Under mediated commits and an arbiter detection floor $δ_M$, it bounds the expected number of unauthorized irreversible actions by $B+m-1+m/δ_M$, a constant in $N$, of which the $B+m-1$ term is decided by a model-free rule and therefore survives a fully colluding verifier. We give a complete evaluation protocol on native Agent-SafetyBench tasks with paired clean and attacked episodes, an outer-state attack suite whose decisive evidence exists only across iterations, per-module ablations, and an adaptive white-box red team.
Problem

Research questions and friction points this paper is trying to address.

safety state
autonomous loop
large language model agents
cross-iteration state
composition failure
Innovation

Methods, ideas, or system contributions that make the work stand out.

LoopHarness
non-decaying safety state
cross-iteration security
autonomous LLM agents
safety composition
💼 Related Jobs
No related jobs found.
Chenhao Wu
Chenhao Wu
University of Chinese Academy of Sciences
H
Haoxuan Jia
Nanyang Technological University
Y
Yang Liu
Supply Chain Tech Team Y, JD.com
Yingguang Yang
Yingguang Yang
University of Science and Technology of China
Y
Yuhan Lin
Fudan University
Chongyang Zhang
Chongyang Zhang
Professor, Shanghai Jiao Tong University
Computer VisionMachine Learningand Artificial Intelligence
H
Hao Zheng
Fullive-AI
Y
Yulin Huang
Supply Chain Tech Team Y, JD.com
J
Jianshen Zhang
Supply Chain Tech Team Y, JD.com
Y
Yongzhi Qi
Supply Chain Tech Team Y, JD.com
S
Shang Luo
Peking University
K
Kefu Xu
Peking University
J
Jifeng Zhu
Fullive-AI
B
Bin Chong
Peking University