StereoPatch: Patch-Aligned RGB-Depth Fusion for Spatial Perception in Robot Manipulation

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为了解决机器人操作中视觉相似但需不同动作的问题,提出StereoPatch方法,通过将深度信息与RGB特征对齐融合来改善空间感知和控制效果。
📝 Abstract
Recent advances in robot imitation learning have produced visuomotor policies that predict actions directly from visual observations. Yet visually similar scenes can require different actions as target position, object height, or contact geometry changes. Pretrained RGB features may map these geometrically distinct states to similar policy inputs, while simply adding depth requires the policy to learn RGB-depth correspondence from the same limited demonstrations used to learn control. We introduce StereoPatch, a patch-aligned RGB-depth representation that binds registered metric geometry directly to the RGB patches used for action prediction. On a shared 2-D patch grid, asymmetric cross-attention incorporates depth information into the corresponding RGB features before action decoding. The resulting StereoPatch Tokens provide a geometry-aware visual representation that can condition general visuomotor policies without changing their underlying learning objectives. Across six real-robot tasks, StereoPatch achieves higher closed-loop success than appearance-only, geometry-only, raw RGB-D, and late-fusion baselines. Additional experiments across three simulation suites evaluate compatibility across visuomotor policy architectures, spatial generalization, and operating limits. Results suggest that resolving control-relevant geometric ambiguity benefits from aligning depth directly with the visual features used for action prediction, rather than supplying it as an independent modality. Project page: https://aus.bot/research/stereopatch/.
Problem

Research questions and friction points this paper is trying to address.

StereoPatch
Robot Manipulation
RGB-Depth Fusion
Spatial Perception
Visuomotor Policies
Innovation

Methods, ideas, or system contributions that make the work stand out.

patch-aligned RGB-depth fusion
asymmetric cross-attention
geometry-aware visual representation
visuomotor policies
Y
Yanan Zhou
School of Computer Science, The University of Sydney, Australia
Z
Zhaoyan Qian
School of Computer Science, The University of Sydney, Australia
J
James Zhao
School of Computer Science, The University of Sydney, Australia
W
Weiming Zhi
School of Computer Science, Australian Centre for Robotics, The University of Sydney, Australia