GuardianBench: A Same-Scene Instruction-Contrastive Benchmark for Latent Contextual Risk in Embodied AI

📅 2026-08-22
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为了解决具身AI中潜在的上下文风险问题,研究引入了GuardianBench基准,通过对比同一场景下的安全/不安全指令对来评估模型性能,并提出了一种轻量级方法提高模型判断准确性。
📝 Abstract
In embodied AI, safety risk can be latent: a benign instruction and a safe scene become hazardous only when composed. Prior work has advanced embodied safety by varying visual contexts or evaluating execution-time dynamics, but the complementary axis of fixing the scene and varying only the instruction remains underexplored. We introduce GuardianBench, an instruction-contrastive benchmark grounded in international safety standards that isolates this latent contextual risk through 3,024 instruction-scene examples organized as same-scene Safe/Unsafe contrastive pairs across various hazard categories. Benchmarking state-of-the-art vision-language models (VLMs) reveals instruction-insensitive verdicts: models disproportionately approve both instructions under a given scene; across the primary models, average pair accuracy is only 24.1%. Our systematic rationale audit localizes the dominant failure: models fail to bind the instruction-relevant cues that differentiate safe from unsafe compositions. As a post-training case study, Verdict Log-Odds Supervision (VLOS), a lightweight verdict-level objective, substantially improves performance on open-weight backbones. Together, our latent contextual risk task formulation, standards-grounded contrastive benchmark construction, pair-level and rationale-level failure diagnosis, and benchmark-enabled verdict calibration establish GuardianBench as a controlled evaluation suite for exposing and improving safety reasoning over instruction-scene compositions under latent contextual risk.
Problem

Research questions and friction points this paper is trying to address.

embodied AI
latent contextual risk
instruction-contrastive
Innovation

Methods, ideas, or system contributions that make the work stand out.

instruction-contrastive benchmark
latent contextual risk
Verdict Log-Odds Supervision (VLOS)
safety reasoning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zhesheng Zhang
Shandong University
Jiahao Lu
Jiahao Lu
National University of Singapore
AGI risksAI controlAI safetyAI alignmentAI security
W
Wei Liu
Shandong University
C
Cong Pan
Nanjing University of Aeronautics and Astronautics
J
Jianhua Yang
Institute of Automation, Chinese Academy of Sciences
Y
Yixiang Chen
Institute of Automation, Chinese Academy of Sciences
H
Hongyuan Yu
Xiaomi Corporation
Mengqi Zhang
Mengqi Zhang
Shandong University
Large Language ModelsData MiningKnowledge Representation Learning
K
Kailin Lyu
Institute of Automation, Chinese Academy of Sciences
Zhumin Chen
Zhumin Chen
Shandong University
Keji He
Keji He
SDU << CASIA & NUS
Cross-modal LearningEmbodied AI