The Guard That Cried Wolf: How Scary Words Make Agent Guardrails Refuse Legitimate Actions

📅 2026-08-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究解决了代理护栏过度拒绝合法操作的问题,通过构建Cautious Bench基准来评估不同设计下的护栏性能,揭示了对象名称对决策的影响。
📝 Abstract
Agent guardrails are checks that approve or refuse each action before an LLM executes it. Sometimes they refuse requests that are genuinely safe. This over-safety blocks deployment when a guardrail refuses an authorized task. Evaluating over-safety is hard: at the boundary an authorized action resembles an unauthorized one, and the safe-versus-unsafe label is a choice of authorization policy, not fixed by the action alone. We argue it therefore requires a benchmark that does not yet exist, one that maps the decision boundary of an ideal guardrail. Harvesting such a benchmark from real data is impractical: boundary cases are hard to collect, their labels hard to verify. The gap is real, so we construct Cautious Bench, the first benchmark to make over-safety the construct for agent guardrails; it codesigns each sample and its label with a stated authorization policy. A build-time gate re-derives every example to certify it, so each label is a mechanical consequence of the policy rather than an annotator's per-sample verdict, a reference against which researchers can measure real guardrails. The benchmark renders 756 Decidable benign/twin pairs, each under three object-name types (2,268 measured pairs), and 40 Undecidable pairs reported separately. Measuring six guardrails from five designs, we find a name-superstition effect: each over-refuses an authorized action more often under a scary-looking object name than a benign one. Since only the object name varies in the aforementioned contrast experiments, the deviation is the name's doing: the guardrails read the surface label, not the authorization context.
Problem

Research questions and friction points this paper is trying to address.

Agent Guardrails
Over-Safety
Authorization Policy
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cautious Bench
over-safety
guardrails
name-superstition effect
Y
Yingjie Zhang
Institute of Information Engineering, Chinese Academy of Sciences, China; School of Cyber Security, University of Chinese Academy of Sciences, China
Y
Yuanbo Xie
Institute of Information Engineering, Chinese Academy of Sciences, China; School of Cyber Security, University of Chinese Academy of Sciences, China
Kai Chen
Kai Chen
Institute of Information Engineering, Chinese Academy of Sciences
Software analysis and testingartificial intelligencesmartphonesprivacy