Beyond Placement and Articulation: Usage-Driven Code Scenes for Embodied Interaction

📅 2026-08-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该研究针对室内场景功能使用建模不足的问题,提出RoomWright框架,通过代码生成3D场景,实现基于使用的物体推理及多部分交互。
📝 Abstract
Indoor scene synthesis provides essential environments for embodied AI, robotic manipulation, and simulation-based policy learning. Recent code-based scene generation methods produce editable and extensible environments, yet they remain focused on visual construction and object-level articulation, leaving the functional usage of scenes largely unmodeled. To address this problem, we present RoomWright, an agentic usage-driven framework for generating 3D scenes represented entirely as code for embodied interaction. RoomWright performs usage-driven object reasoning, which treats each anchor as a task centre and admits task-required objects and their affordances. A code agent further enables multi-part interaction by compiling each interaction into a trigger, condition, effect rule that updates structured object states, capturing causal dependencies across objects. Moreover, since manipuland orientation is ambiguous and hard to recover from pixels, RoomWright alleviates this via annotation-informed usage-guided orientation. Extensive experiments demonstrate the effectiveness of our method. The resulting scenes are executable, editable, and simulation-ready, providing interactive environments for embodied AI and policy learning.
Problem

Research questions and friction points this paper is trying to address.

indoor scene synthesis
embodied interaction
functional usage
code-based generation
object-level articulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

usage-driven object reasoning
code agent
embodied interaction
task-based generation
annotation-informed orientation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zijian Xiao
Shanghai Jiao Tong University
Zipeng Ye
Zipeng Ye
Tsinghua University
Computer GraphicsArtificial Intelligence
J
Jinkun Hao
Shanghai Jiao Tong University
X
Xiong Yang
Shanghai Jiao Tong University
Y
Yuchen Xie
Meituan
Ran Yi
Ran Yi
Associate Professor, Shanghai Jiao Tong University
Computer VisionComputer Graphics