NeoWorld-Pro: Programming Interactive Scenes from Monocular Images for Embodied Simulation

📅 2026-08-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为了解决单目图像转换为可模拟场景的难题,NeoWorld-Pro利用MLLMs将图像转化为包含物体几何、关节和物理属性的程序,并通过物理引擎迭代优化。
📝 Abstract
The advancement of Embodied AI necessitates high-quality simulation assets that faithfully mirror the real world. However, transforming raw visual observations into simulation-ready scenes remains challenging due to the lack of physical grounding and scene-level interactivity in current image-to-URDF methods. We propose NeoWorld-Pro, a framework that reformulates monocular scene reconstruction as procedural programming for interactive 3D environments. Leveraging the zero-shot reasoning and code synthesis capabilities of MLLMs, NeoWorld-Pro converts a single RGB image into executable programs specifying object geometry, articulation, and physical properties. A physics-in-the-loop mechanism then iteratively refines the generated programs by validating their execution in a physics engine, enforcing physically plausible articulations, valid object compositions and interactions, and accurate spatial relationships. Experiments show that NeoWorld-Pro outperforms open-loop and prior monocular reconstruction methods, while enabling complex downstream tasks such as stable stacking and fine-grained manipulation.
Problem

Research questions and friction points this paper is trying to address.

Embodied AI
monocular images
scene reconstruction
physical grounding
scene-level interactivity
Innovation

Methods, ideas, or system contributions that make the work stand out.

procedural programming
zero-shot reasoning
code synthesis
physics-in-the-loop
🔎 Similar Papers
No similar papers found.
Y
Yumeng He
Shanghai Jiao Tong University
Y
Yichen Song
Shanghai Jiao Tong University
X
Xiaotian Yang
Huazhong University of Science and Technology
W
Weijia Zhang
Shanghai Jiao Tong University
Z
Zanwei Zhou
Shanghai Jiao Tong University
J
Junru Gong
Shanghai Jiao Tong University
X
Xiaokang Yang
Shanghai Jiao Tong University
Yunbo Wang
Yunbo Wang
Associate Professor, Shanghai Jiao Tong University
Machine LearningComputer Vision