Pointing-VLA: Typed Spatial Grounding Interfaces for Vision-Language-Action Manipulation

📅 2026-08-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决VLA模型中空间定位接口脆弱问题,提出Pointing-VLA方法,通过几何特定头部预测点、热图和轨迹,直接提供空间目标,提高机器人执行效率。
📝 Abstract
Vision-language-action (VLA) models often expose spatial grounding through autoregressive text coordinates or opaque action tokens, creating brittle interfaces between multimodal reasoning and robot execution. We present Pointing-VLA, a typed hidden-state spatial readout built on Embodied-R1. Geometry-specific heads predict normalized points, object-functional grounding (OFG) heatmaps, and visual trajectories without serializing geometry as text. For the evaluated Bridge/WidowX and physical pick-place deployments, an explicit execution contract assigns PICK to source-conditioned OFG and PLACE to Pointing, providing direct stage-aligned spatial targets. Pointing-VLA achieves SOTA performance on Bridge/WidowX, averaging 72.9\% across the evaluated four-task set without Bridge-specific finetuning under collision-enabled CuRobo execution. Pointing and OFG show complementary strengths across native and cross-dataset evaluations. The OFG/contact readout transfers to NORA-1.5, preserving or improving success while reducing recorded controller time by more than 20$\times$; typed heads are also 6.68--6.90$\times$ faster than Embodied-R1 text decoding on a shared external suite. When integrated as spatial guidance for a $π_{0.5}$ action policy, Pointing-VLA raises autonomous real-robot success from 52.7\% to 80.7\% across three visual contexts. These results establish typed spatial readouts as an efficient, inspectable interface between embodied reasoning and robot execution.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action
Spatial Grounding
Multimodal Reasoning
Robot Execution
Innovation

Methods, ideas, or system contributions that make the work stand out.

Typed Spatial Readout
Object-Functional Grounding (OFG)
Visual Trajectories
Execution Contract
Spatial Guidance
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Xiwen Chen
Xiwen Chen
Clemson University
Deep LearningMultimodalComputer VisionTime Series AnalysisVLM/LLM
Z
Zelin Li
SIGS, Tsinghua University, China; Wuhan University of Technology, China
Z
Zhiruo Zhou
SIGS, Tsinghua University, China; Wuhan University of Technology, China
H
Huiming Chen
City University of Hong Kong, Hong Kong (China SAR)
C
Chenwei Wang
AiDlab, Hong Kong (China SAR)
Xiaojun Zhu
Xiaojun Zhu
SIGS, Tsinghua University, China