RefGuard: Identity-Aware Language-Guided Robot Manipulation via Joint Target-Anchor-Frame Grounding

📅 2026-09-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决机器人在复杂场景中因对象识别错误导致的操作失误,提出RefGuard框架,通过联合目标-锚点-参照系定位方法提高操作准确性。
📝 Abstract
Vision-language-action (VLA) models have substantially advanced language-guided robot manipulation, yet reliable execution still hinges on identifying which physical object an instruction refers to. In cluttered scenes containing repeated objects, ambiguous anchors, or frame-dependent spatial terms, a robot can execute a geometrically valid action on a semantically compatible but unintended instance; we call this failure an identity switch. The referent is jointly determined by three coupled latent variables: the target, the anchor, and the reference frame, so committing to any one of them before execution turns residual ambiguity into a silent and irreversible error. We propose RefGuard, an identity-aware grounding framework that delays commitment by maintaining a joint posterior over all three variables. RefGuard builds a frame-conditioned object-centric scene graph from RGB-D observations, separating frame-independent geometry from directional relations, and routes the posterior through a decision policy that executes, clarifies, reobserves, or aborts. On a real UF850 arm, RefGuard records no identity switch on any ambiguity-stress trial and executes correctly on 90.0% of them, whereas fine-tuned VLA and LLM (Large Language Model)-based baselines switch identity in 33-46% of the same trials, while retaining 93.3% success on unambiguous scenes and recovering from post-grounding scene changes in 86.7% of trials. On a 3200-episode procedural suite, it raises correct execution on solvable instructions from 56.6% to 80.5% over the ablation that commits to the anchor and frame before the target, while deferring less often (19.5% vs. 43.4%).
Problem

Research questions and friction points this paper is trying to address.

Identity Switch
Robot Manipulation
Vision-Language-Action Models
Ambiguity
Cluttered Scenes
Innovation

Methods, ideas, or system contributions that make the work stand out.

identity-aware grounding
joint posterior
frame-conditioned object-centric scene graph
L
Lan Wei
Imperial College London, UK
K
Kangyi Lu
Imperial College London, UK
Y
Yongchen Wang
Imperial College London, UK
C
Chenmeng Bi
Imperial College London, UK
Q
Qi Chen
Imperial College London, UK
Hanlin Niu
Hanlin Niu
RACE, United Kingdom Atomic Energy Authority, UK and Oxford Robotics Institute, University of Oxford, UK
Yip Fun Yeung
Yip Fun Yeung
PhD, MIT
RoboticsControlAI
Dandan Zhang
Dandan Zhang
Imperial College London
RoboticsAI