What Matters, When? Diagnosing and Improving Conditional Visual Grounding in Visuomotor Imitation Policies

📅 2026-09-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过引入干扰物和使用ACT方法,诊断并改进了视觉运动模仿策略中的条件视觉定位问题,提高了目标选择的鲁棒性。
📝 Abstract
Visuomotor imitation policies can achieve high performance under in-distribution visual conditions yet fail when visually similar objects or receptacles are introduced. We study this behavior as a problem of conditional visual grounding: the visual target required for successful control changes with the manipulation phase and, in more complex tasks, with the observed task state. Using Action Chunking with Transformers (ACT), we systematically introduce distractor objects and receptacles with controlled color and shape similarity and localize failures to picking and placement. We find that distractor sensitivity is specific to both the type of visual similarity and the manipulation stage. Guided by this diagnosis, we evaluate distractor augmentation, phase-dependent attention regularization, and appearance-based visual prompting as complementary interventions for improving target selection while preserving spatial information required for control. These interventions substantially improve robustness in simulation and on a physical UR3e. We further examine the same failure pattern in a pretrained vision-language-action policy on a state-conditioned instrument-handling task, where the observed state of a medical instrument determines the correct destination. Together, the results show that visual distractors can cause incorrect object or destination selection even when the underlying manipulation skill remains intact, and that explicitly improving target selection can substantially recover performance across distinct visuomotor policy-learning regimes.
Problem

Research questions and friction points this paper is trying to address.

visuomotor imitation policies
visual grounding
distractor objects
manipulation phase
task state
Innovation

Methods, ideas, or system contributions that make the work stand out.

conditional visual grounding
Action Chunking with Transformers (ACT)
distractor augmentation
phase-dependent attention regularization
appearance-based visual prompting
💼 Related Jobs
No related jobs found.
V
Vivek Chavan
Fraunhofer Institute for Production Systems and Design Technology IPK
Pengtao Xie
Pengtao Xie
Associate Professor, UC San Diego; Adjunct Faculty, MBZUAI
Machine Learning
Y
Yahuan Shi
Technische Universität Berlin
Oliver Heimann
Oliver Heimann
Fraunhofer Institute for Production Systems and Design Technology
Robotics and Computer Vision
K
Kevin Haninger
Fraunhofer Institute for Production Systems and Design Technology IPK
Jörg Krüger
Jörg Krüger
Professor Industrial Automation Technology, TU Berlin
AutomationRoboticsComputer VisionControl