Action-Sketcher: From Reasoning to Action via Visual Sketches for Long-Horizon Robotic Manipulation
This work addresses the challenges of spatial reference ambiguity, task decomposition difficulty, and opaque decision-making that arise when using natural language instructions for long-horizon robotic manipulation. To overcome these issues, the authors propose employing editable visual sketches as an explicit intermediate representation, establishing a closed-loop “perceive–reason–sketch–act” workflow that precisely maps linguistic intent onto scene geometry. Through a multi-stage curriculum training framework, the approach integrates modality alignment, language-to-sketch consistency constraints, and sketch-guided reinforcement imitation learning to enable adaptive policy learning and real-time action prediction. Evaluated in both simulated and real-world complex environments, the method significantly improves task success rates and dynamic robustness while supporting human-in-the-loop correction and yielding interpretable, intervenable decision processes.