Self Supervised Learning from Automatically Generated Demonstrations for Visual Robotic Manipulation
This work proposes a self-supervised visual manipulation method that eliminates the need for manual programming, human demonstrations, or extrinsic camera-robot calibration, which traditionally incur high deployment costs. The robot autonomously generates visual demonstrations near the target pose and learns relative pose corrections directly from wrist-mounted RGB images. A coarse-to-fine two-stage control strategy is employed, and an image-pose paired dataset is constructed using ROS 2 and Isaac Sim. A convolutional neural network regresses relative translation and rotation from single-frame RGB inputs. In simulation, the planar positioning error decreases from 9.69 mm to 5.38 mm. On a real UR5e robot, the method achieves grasping success rates of 66.6% and 63.6% on two object categories and demonstrates robustness under rotational perturbations.