ManiSkillFormer: Demonstration-Free Compositional Manipulation via Task-Conditioned Geometric Contracts

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
ManiSkillFormer通过任务条件几何契约和神经符号框架解决无示范复合机器人操作问题,实现高成功率的多功能操作。
📝 Abstract
We present ManiSkillFormer, a neuro-symbolic framework for demonstration-free and compositional robotic manipulation. Instead of learning end-to-end visuomotor policies, ManiSkillFormer introduces task-conditioned geometric contracts that explicitly structure the interface between perception and action. Each manipulation skill declares the semantic geometric primitives required for execution, such as object keypoints and surface normals. Building on human-defined skill structures, LLM agents generate these contracts and corresponding motion templates for different objects and task contexts. These contracts guide the perception module to ground task-relevant 3D primitives from observations, which are then used to instantiate reusable motion templates stored in a skill library. We evaluate ManiSkillFormer on Galaxea R1-Lite dual-arm robot across three settings: zero-shot pick-and-place over 8 object categories with 30 different instances, functional manipulation tasks including unscrewing, pouring, pressing, and folding, and 3 long-horizon tasks. ManiSkillFormer achieves higher average success rates than the evaluated baselines and two ablated pipelines: 88.24% for demonstration-free pick-and-place, 75.00% average success on functional manipulation and 50--80% completion rates across the long-horizon tasks. These results show that our design enables composable and reusable manipulation across objects and tasks without per-object policy fine-tuning or additional robot demonstrations.
Problem

Research questions and friction points this paper is trying to address.

Demonstration-Free
Compositional Manipulation
Task-Conditioned
Geometric Contracts
Robotic Skills
Innovation

Methods, ideas, or system contributions that make the work stand out.

task-conditioned geometric contracts
neuro-symbolic framework
compositional manipulation
demonstration-free
reusable motion templates
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
P
Peiqi Yu
Carnegie Mellon University, Pittsburgh, PA 15213, USA
Mosam Dabhi
Mosam Dabhi
PhD Student, Carnegie Mellon University
Self-supervisionAuto-labeling3D reconstructionMulti-view GeometryStructured optimization
S
Shangtao Li
Carnegie Mellon University, Pittsburgh, PA 15213, USA
B
Bowei Li
Carnegie Mellon University, Pittsburgh, PA 15213, USA
L
Laszlo Jeni
Carnegie Mellon University, Pittsburgh, PA 15213, USA
Changliu Liu
Changliu Liu
Associate Professor, Carnegie Mellon University
Roboticshuman-robot interactionsmotion planningoptimizationmulti-agent system