Self Supervised Learning from Automatically Generated Demonstrations for Visual Robotic Manipulation

📅 2026-08-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work proposes a self-supervised visual manipulation method that eliminates the need for manual programming, human demonstrations, or extrinsic camera-robot calibration, which traditionally incur high deployment costs. The robot autonomously generates visual demonstrations near the target pose and learns relative pose corrections directly from wrist-mounted RGB images. A coarse-to-fine two-stage control strategy is employed, and an image-pose paired dataset is constructed using ROS 2 and Isaac Sim. A convolutional neural network regresses relative translation and rotation from single-frame RGB inputs. In simulation, the planar positioning error decreases from 9.69 mm to 5.38 mm. On a real UR5e robot, the method achieves grasping success rates of 66.6% and 63.6% on two object categories and demonstrates robustness under rotational perturbations.
📝 Abstract
Robotic manipulation often requires object specific programming, manual data annotation, or calibrated perception pipelines, which limits rapid deployment in practical settings. Learning from demonstration offers a more direct alternative, but collecting demonstrations can still demand human teleoperation or kinesthetic teaching. This paper presents a self supervised visual manipulation method in which a robot automatically generates demonstrations around a target pose and learns relative pose corrections directly from wrist mounted RGB images. The proposed pipeline uses ROS~2 and Isaac Sim to collect labeled image-pose pairs without requiring explicit camera to robot extrinsic calibration. Separate datasets are generated for planar refinement and coarse three dimensional approach, and a convolutional network is trained to regress relative translation and rotation from single frame RGB observations. During execution, a coarse to fine controller first approaches the object using models trained with height variation and then refines the final alignment using planar data. The method is evaluated both in simulation and on a real UR5e collaborative robot equipped with a gripper and a monocular camera. In simulation, the refinement stage reduces the final planar dispersion from 9.69 mm to 5.38 mm. In real world experiments, the system performs end to end grasp attempts on three physical objects and reaches success rates of 66.6% and 63.6% for two objects without object rotation, while still maintaining partial robustness under rotated conditions. These results show that automatically generated demonstrations can support practical visual manipulation with limited setup effort, while also exposing remaining challenges in depth prediction and object dependent generalization.
Problem

Research questions and friction points this paper is trying to address.

robotic manipulation
learning from demonstration
self-supervised learning
visual perception
automatic data generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

self-supervised learning
automatically generated demonstrations
visual robotic manipulation
pose regression
coarse-to-fine control
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
A
Andres Rivas
Robotics and AI Lab, Technological University of Uruguay, Rivera, Uruguay
A
Anselmo R. Cukla
Center of Technology, Federal University of Santa Maria, Santa Maria, Brazil
R
Rodrigo S. Guerra
Computer Science Center, Federal University of Rio Grande, Rio Grande, Brazil
B
Bruna V. Guterres
Robotics and AI Lab, Technological University of Uruguay, Rivera, Uruguay
R
Ricardo B. Grando
Robotics and AI Lab, Technological University of Uruguay, Rivera, Uruguay