Contact-Anchored Policies: Contact Conditioning Creates Strong Robot Utility Models
Language instructions are often too abstract to support robust robotic manipulation in complex physical interactions. To address this limitation, this work proposes replacing linguistic commands with spatial contact points as the conditioning signal for policy learning, constructing a modular utility model library, and integrating it with EgoGym—a lightweight simulation platform—to enable rapid real-to-sim closed-loop iteration. Using only 23 hours of demonstration data, the approach achieves out-of-the-box, zero-shot generalization across environments and robot embodiments on three fundamental manipulation tasks, outperforming state-of-the-art vision-language-action models by 56% in performance. The core innovation lies in the novel contact-point-conditioned policy architecture, which substantially enhances both generalization capability and deployment efficiency.