IM-ENGINE: Image Editing for Embodied Data Generation

📅 2026-09-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
IM-ENGINE通过图像编辑和模拟器先验解决学习式操作中语义与物理执行监督不足的问题,生成可用于机器人学习的执行数据。
📝 Abstract
Learning-based manipulation requires supervision that is both semantically meaningful and physically executable, but current data pipelines often provide only one of these properties. Human demonstrations capture intent but are costly to collect and constrained by the human-robot embodiment gap, while simulation can scale data generation but often under-specifies functional behavior. We present IM-ENGINE, a simulator-grounded pipeline that uses image editing as an intermediate representation for embodied data generation. Given a rendered scene with known geometry, depth, segmentation, and camera parameters, IM-ENGINE edits the image to inject task-relevant semantics, recovers explicit 3D state using simulator priors and an unchanged anchor object, refines the state in physics, and converts it into robot-executable supervision. We instantiate the pipeline for dexterous grasp synthesis and goal-state generation. For grasping, IM-ENGINE generates a human grasp in image space, recovers the hand-object interaction, retargets it to a robot hand, and refines it into physically validated robot grasps. For goal generation, it edits a rendered scene into a desired outcome, recovers the target-object pose, and refines it into physically valid, semantically meaningful goals and trajectories. This combination of generative semantic priors and simulator grounding enables scalable task-relevant supervision for robot learning.
Problem

Research questions and friction points this paper is trying to address.

learning-based manipulation
semantically meaningful
physically executable
human-robot embodiment gap
simulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

image editing
simulator priors
physically executable supervision
embodied data generation
task-relevant semantics