Acting in Meters: Learning Metric Interactions for Precise Robotic Manipulation

📅 2026-09-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出了一种度量交互框架,通过显式建模对象和场景级别的物理空间交互来提高机器人操作精度,从而改进了视觉-语言-动作模型和世界-动作模型。
📝 Abstract
Vision-Language-Action models and World-Action Models have advanced language-conditioned robotic manipulation, yet often leave metric relations among actions, objects, and scene geometry implicit. Human manipulation combines semantic understanding of task-relevant objects with spatial feedback that guides hand motion relative to objects and their surroundings. Inspired by this, we introduce a metric interaction framework that models object-level and scene-level interactions in physical Cartesian space at a shared metric scale. At the object level, Interaction-Centric Tokens (ICTs) explicitly represent end-effector pose trajectories relative to manipulated objects and are jointly denoised with actions, providing physically grounded interaction supervision. At the scene level, the Metric Action Interaction Field (MAIF) uses action and ICT queries to attend to metric scene point-cloud features and learns geometry-conditioned action corrections. Through two-stage adaptation, our framework improves diverse VLA and WAM baselines with a small number of additional parameters and training steps. Experiments demonstrate average success-rate gains of 0.80 and 3.59 percentage points on LIBERO and RoboTwin~2.0, respectively, alongside gains of 6.80 percentage points on real-world tasks and 7.45 percentage points on their out-of-distribution variants.
Problem

Research questions and friction points this paper is trying to address.

Metric Interactions
Robotic Manipulation
Vision-Language-Action Models
World-Action Models
Spatial Feedback
Innovation

Methods, ideas, or system contributions that make the work stand out.

Metric Interaction Framework
Interaction-Centric Tokens (ICTs)
Metric Action Interaction Field (MAIF)
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Lijie Wang
Lijie Wang
FASTLab, Zhejiang University
Multi-Robot SystemMulti-Sensor FusionRobot Perception
Z
Zheng Lu
Peking University
Yiming Wang
Yiming Wang
Shanghai Jiao Tong University
Large Language ModelsComplex ReasoningAI Interpretability
H
Heyang Yu
Tsinghua University
K
Kenghou Hoi
Zhejiang University
B
Bowen Hu
Zhejiang University
Di Cui
Di Cui
Xidian University
Software ArchitectureSelf-adaptive SystemsRefactoringSoftware Performance
T
Tianyu Xin
Tsinghua University
H
Haoran Liao
Sun Yat-sen University
W
Wanqi Zhong
Tsinghua University
X
Xingjie Fan
Tsinghua University
Y
Yizhao Xu
Peking University
Z
Ziliang Wang
Tsinghua University
Fei Gao
Fei Gao
Associate Professor, Zhejiang University
Aerial RoboticsMotion PlanningAutonomous Navigation
Y
Yiming Li
Tsinghua University