GIF: Agentic Generation of Interactive and Functional Object Compositions for Robot Learning

📅 2026-09-05
📈 Citations: 0
Influential: 0
📄 PDF
📝 Abstract
Robot manipulation foundation models require scalable evaluation and data generation across diverse scenarios, with simulation providing an environment for both. Automated scene generation offers a promising path, yet prior work has largely emphasized coarse-grained scene layouts rather than fine-grained functional object compositions. Motivated by this gap, we present GIF, an agentic Generation framework for Interactive and Functional object compositions. In this framework, we recast this problem as disentangled reconstruction followed by relative pose recovery. CoGen produces instance-disentangled meshes with coarse initial poses leveraging complementary strengths of 2D and 3D generative models. GPRM refines the relative pose under joint geometric and physical guidance, and a VLM verifier selects the candidate that best matches the structured specification. We further construct a benchmark spanning eight representative contact-geometry classes and compare with state-of-the-art generators; GIF improves both asset quality and relation matching, while reducing collision rate to below 1%. Finally, we synthesize data for policy learning, revealing diversity scaling in both simulation and real-world deployment.
Problem

Research questions and friction points this paper is trying to address.

robot manipulation
scene generation
functional object compositions
Innovation

Methods, ideas, or system contributions that make the work stand out.

disentangled reconstruction
relative pose recovery
interactive and functional object compositions
💼 Related Jobs
No related jobs found.
Long Xu
Long Xu
Ningbo University, Peng Cheng Laboratory
image/signal processingvideo codingespecially rate control of video codingimage/signal
Z
Zhiqi Zhang
Galbot; Peking University
M
Mi Yan
CFCS, School of CS, Peking University; Galbot
S
Shengliang Deng
Galbot; The University of Hong Kong
C
Chong Xia
Galbot; Tsinghua University
M
Mingyu Dong
Galbot; Tsinghua University
Jiayi Chen
Jiayi Chen
Peking University
Robotics3D Vision
J
Jiangran Lyu
CFCS, School of CS, Peking University; Galbot
Fei Gao
Fei Gao
Associate Professor, Zhejiang University
Aerial RoboticsMotion PlanningAutonomous Navigation
Z
Zhizheng Zhang
Galbot; Beijing Academy of Artificial Intelligence
H
He Wang
CFCS, School of CS, Peking University; Beijing Academy of Artificial Intelligence