Affordance-Based Manipulation Planning with Text Goals and Sim-to-Real Generalisation via Real-to-Sim Image Conversion
This work addresses the challenge of enabling robots to reliably plan and execute tasks from natural language instructions in complex, dynamic environments with occlusions, while achieving effective sim-to-real transfer. To this end, the authors propose a planning framework that integrates functional affordance recognition with visual action-effect prediction, leveraging visual forward reasoning to anticipate future states. A multimodal text-image matching module is introduced to evaluate the consistency between candidate action sequences and the linguistic goal. Furthermore, a real-to-sim image stylization mechanism is designed to enhance perceptual robustness in real-world settings. Experimental results demonstrate that the proposed approach successfully accomplishes challenging manipulation tasks on both simulated and physical robot platforms, significantly improving language-conditioned generalization from simulation to reality.