MANIGUARD: A Benchmark and Data Suite for Specification-Grounded Safety Evaluation and Improvement of Robotic Manipulation

📅 2026-08-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决机器人操作安全性评估不足的问题,本文提出ManiGuard框架,通过安全注释轨迹生成管道和任务套件来独立于任务成功进行安全性评估和改进。
📝 Abstract
Foundation-model policies for robotic manipulation are advancing rapidly on task success, but rigorous evaluation of whether they succeed safely is still lacking. We introduce ManiGuard, a specification-grounded framework for evaluating and improving the safety of foundation-model manipulation, comprising the ManiGuard-Bench task suite and a paired safety-annotated trajectory-generation pipeline. ManiGuard-Bench organizes six contact-rich household task families into 200 locked base tasks along a skill $\times$ constraint taxonomy, with safety specified independently of task success. Each task is evaluated under one in-distribution and four single-axis out-of-distribution perturbations that hold the safety specification fixed, giving 1,000 locked scenarios. Every rollout is runtime-checked by LTL$_f$-grounded automaton monitors over physics-grounded predicates rather than learned classifiers or LLM judges, in simulation and on a physical Franka platform. The pipeline pairs an automated motion-planning generator with human teleoperation, annotated by the same per-step monitor, and directly supports safety-aware fine-tuning; we release 8,000 safety-annotated demonstrations, 40 per base task. Benchmarking zero-shot and fine-tuned VLAs across more than 23,000 rollouts, we find: (i) safety must be evaluated independently of task success, as 6-21% of successful rollouts violate the specification; (ii) fine-tuning on our suite raises safe task completion from near zero to 7.5-29.8% and engaged-and-safe behavior from 16-40% to 51-72%; but (iii) a gap remains that scaling demonstrations does not close, with 21-42% of engaged rollouts still violating, two of six families below 2% safe success for every policy, and these failures persisting under distribution shift and on hardware.
Problem

Research questions and friction points this paper is trying to address.

robotic manipulation
safety evaluation
foundation-model policies
Innovation

Methods, ideas, or system contributions that make the work stand out.

specification-grounded
safety-annotated
automaton monitors
fine-tuning
out-of-distribution
Y
Yiyan Peng
Northwestern University
P
Philip Wang
Northwestern University
S
Simon Sinong Zhan
Northwestern University
Y
Yiqi Lyu
Northwestern University
Z
Zhenyang Ni
Northwestern University
J
Jixin Yan
Northwestern University
F
Fiorelli Wong
Northwestern University
Ruochen Jiao
Ruochen Jiao
Northwestern University
Trustworthy Machine LearningEmbodied AILarge Language Models
H
Hang Yin
Stanford University
X
Xinyu Cao
Northwestern University
Huajie Shao
Huajie Shao
Assistant Professor, William and Mary
Physcis-guided Machine LearningAI for Cyber-physical SystemsIOT
Manling Li
Manling Li
Assistant Professor at Northwestern University
Natural Language ProcessingVision-LanguageEmbodied Agents
Ruohan Zhang
Ruohan Zhang
Stanford University
RoboticsCognitive ScienceBrain-Machine InterfaceMachine LearningArt
Q
Qi Zhu
Northwestern University