SWIM: Vision-Language-Grounded Soft Whole-Body Interactive Manipulation

📅 2026-09-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究提出SWIM框架,通过结合视觉-语言动作策略和软体机器人特性,解决从语言和视觉信息到执行命令序列的转换问题。
📝 Abstract
Soft and continuum robots enable manipulation through distributed body deformation and contact, yet translating language and visual context into executable whole-body actuation remains a fundamental challenge. We present SWIM, a framework that maps an initial RGB observation and a language instruction to a complete actuation-command sequence. Its vision-language-action (VLA) policy, SWIM-VLA, combines a diffusion action head with Visual Soft Proprioception (VSP) through a shared representation of RGB observations, language instructions, and tendon states. The diffusion head models conditional distributions of expert command chunks, while VSP supervises ordered body-anchor predictions using simulation ground truth, encouraging the representation to retain body geometry when learning from limited demonstrations. Embodied mechanical intelligence supports physical execution of command sequences generated through iterative virtual rollout from evolving simulated observations, with intrinsic compliance providing local contact adaptation without online policy queries. We evaluate SWIM on packing, reaching, and grasping on a planar tendon-driven soft robot, with grasping targets anchored. In simulation, SWIM-VLA achieves success rates of 100\%, 96\%, and 88\%, respectively, outperforming an adapted OpenVLA-OFT baseline and controlled ablations. On hardware, SWIM achieves success rates of 100\%, 80\%, and 75\%, compared with 75\%, 40\%, and 25\% for direct online deployment of the same policy checkpoint.
Problem

Research questions and friction points this paper is trying to address.

vision-language-grounded
whole-body interaction
soft robots
continuum robots
actuation
Innovation

Methods, ideas, or system contributions that make the work stand out.

diffusion action head
Visual Soft Proprioception (VSP)
vision-language-action (VLA) policy
💼 Related Jobs
No related jobs found.
T
Tingcong Liu
Nanyang Technological University, Singapore
Aye Phyu Phyu Aung
Aye Phyu Phyu Aung
Institute for Infocomm Research (I2R)
Generative ModelsReinforcement Learning
Junjie Xiong
Junjie Xiong
Assistant Professor of Computer Science, Missouri S&T
Network SecuritySoftware SecurityWeb Security
S
Siyi Ma
Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, United Arab Emirates
Bo An
Bo An
Nanyang Technological University
Artificial intelligencemulti-agent systemsgame theoryreinforcement learningoptimization
K
Ke Wu
Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, United Arab Emirates
S
Senthilnath Jayavelu
National University of Singapore, Singapore