DeMoBot: Deformable Mobile Manipulation with Vision-based Sub-goal Retrieval
To address the weak generalization of few-shot (20 demonstrations) imitation learning in partially observable environments, this paper proposes a novel demonstration-driven framework for mobile manipulation tasks. Methodologically, it introduces—first in the literature—a vision foundation model–based (e.g., CLIP) demonstration snippet retrieval mechanism that matches observations via visual similarity; integrates trajectory similarity and forward-reachability constraints to filter feasible subgoals; and employs a goal-conditioned diffusion-based motion policy for action generation. The core contribution lies in abandoning end-to-end fitting in favor of a modular, interpretable, and verifiable architecture that ensures subgoal feasibility and policy robustness. Evaluated on both simulation and real-world Spot robot platforms, the framework achieves significantly higher success rates than state-of-the-art baselines: 85%/80% (sim/real) for gap coverage, 87.5%/70% for tabletop cleanup, and 47.5%/35% for curtain opening.