DeicticVLA: Unifying Instruction Modes Based on Language and Deictic Gestures in a Single VLA

📅 2026-08-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过DeicticVLA模型整合了基于语言和指示手势的三种指令模式,以解决视觉-语言-动作模型中目标识别不准的问题,提高了在新布局下的表现。
📝 Abstract
Vision-Language-Action models (VLAs) allow users to specify manipulation tasks in natural language, but distinguishing a target or placement goal among objects of the same category or similar appearance requires detailed expressions that VLAs may not use reliably. We propose DeicticVLA, which canonicalizes Language Instruction (LI), Vision-Language Instruction (VLI), and Visual Instruction (VI) into a text prompt and deictic masks through text-prompt completion and deictic gesture grounding, enabling a single pretrained VLA to handle all three instruction modes. With a shared backbone, demonstrations, and matched training steps, we compare two RGB visual prompting methods, two separate-channel mask prompting methods, and three training strategies in simulation. Under two-stage training, the four prompting methods achieve high in-distribution success but differ in their ability to use deictic masks in unseen layouts. Across methods, training-strategy ablations show that two-stage training improves such use, while retaining second-stage LI data mitigates forgetting without reducing VLI and VI performance. In three real-world tasks, one policy supports all modes. VLI and VI outperform LI under unseen expressions, appearance changes, and novel objects. For unseen categories, both achieve 100% success, compared with 16.7% for jointly trained LI. These results demonstrate the unified three-mode interface and guide DeicticVLA design.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action
natural language
distinguishing targets
instruction modes
deictic gestures
Innovation

Methods, ideas, or system contributions that make the work stand out.

DeicticVLA
Vision-Language-Action models
deictic masks
instruction modes unification
two-stage training
K
Kango Yanagida
Dept. of Systems Innovation, Graduate School of Engineering Science, The University of Osaka, Japan.
Tatsuya Aoki
Tatsuya Aoki
The University of Osaka
RoboticsArtificial Intelligence
Yuichiro Yoshikawa
Yuichiro Yoshikawa
大阪大学
Takato Horii
Takato Horii
The University of Osaka, Japan
Robotics