🤖 AI Summary
This study addresses the challenges of real-world data scarcity and insufficient generalization in imitation learning for surgical robotics by proposing the SurgVIL framework. This approach innovatively leverages open-source surgical videos to augment training data, utilizing kinematic estimation to provide weak supervision signals that effectively compensate for the absence of ground-truth action labels. Experimental results demonstrate that SurgVIL significantly enhances policy generalization on da Vinci robotic systems, particularly when manipulating real tissues and in out-of-distribution scenarios. By validating the feasibility of integrating multi-source heterogeneous data for weakly supervised learning, this research offers a scalable pathway to overcome critical data bottlenecks in surgical robot autonomy.
📝 Abstract
Learning-based surgical robot autonomy requires large-scale demonstrations with synchronized videos and robot actions, but such data are exceedingly rare in clinical or realistic tissue settings because robot kinematics are typically inaccessible outside controlled research systems. In contrast, phantom data collected on research platforms provide accurate action labels but lack the visual diversity of real tissue. We propose SurgVIL, a framework for scaling surgical robot imitation learning using open-source surgical videos. SurgVIL combines kinematically labeled phantom robot demonstrations with surgical videos from open-source datasets and online sources for policy learning. Since these videos lack robot motion labels, we estimate approximate kinematics as weak supervision. We evaluate SurgVIL on two da Vinci robot tasks: needle pick-up and cholecystectomy cutting. Across ACT, $π_0$, and GR00T-H backbones, adding surgical videos substantially improves generalization to real-tissue and out-of-distribution settings, suggesting a scalable path from phantom training toward generalizable surgical robot policies.