Dreaming the Sound of Contact: Leveraging Video and Audio Generation for Zero-Shot Force-Aware Manipulation and Data Generation

📅 2026-09-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过结合生成的视频和音频信息,为机器人操作任务提供力觉反馈,解决了纯运动轨迹在接触丰富任务中因缺乏力信息而导致失败的问题。
📝 Abstract
Recent advances in video generation allow robots to learn manipulation trajectories from generated videos. However, these approaches produce purely kinematic trajectories that lack force information, causing failures in contact-rich tasks where appropriate contact forces are essential for success. In this work, we explore augmenting generated video with audio to shape a bounded, time-varying desired-force profile using the loudness of generated contact sounds. We present a pipeline that jointly leverages generated video and audio to derive motion trajectories and corresponding desired-force profiles from a structured natural-language task prompt. We execute these force-aware trajectories on a Franka Panda robot using a closed-loop force regulator that tracks the audio-shaped force profile during contact. We evaluate our pipeline on multiple tasks that require making contact and demonstrate successful manipulation where a kinematic-only baseline fails. We also use the pipeline as a data generation engine to train policies that achieve the tasks in a closed-loop manner. Project website, videos, and dataset: https://dreamingcontactsound.github.io/
Problem

Research questions and friction points this paper is trying to address.

video generation
force information
contact-rich tasks
audio generation
manipulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

audio-visual generation
force-aware manipulation
zero-shot learning
closed-loop force control
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.