DesignAgent3D: Interactive 3D Scene Editing via Designer-like Multimodal Reasoning

📅 2026-08-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决3D场景编辑中自然语言请求语义不明确和空间定位不准的问题,提出DesignAgent3D框架,通过计划-感知-行动的模式与用户互动,实现精确修改。
📝 Abstract
Text guided 3D scene editing provides an intuitive interface for modifying reconstructed environments, but remains difficult because natural language design requests are often semantically underspecified and must be grounded in cluttered 3D scenes. Existing methods typically formulate the task as one-shot conditional generation from a single prompt, failing to resolve ambiguous user intents or achieve precise spatial grounding. Consequently, they suffer from severe object localization drift, tracking failure under occlusions, and the notorious multi-view "sticker effect." To overcome these limitations, we present DesignAgent3D, an interactive multimodal agentic framework that reformulates 3D scene editing as a designer-like Plan-Perceive-Act paradigm. The agent first plans by interacting with the user to clarify underspecified design goals, then perceives by grounding the intended edit to specific objects or regions in the 3D scene, and finally acts by applying controlled visual modifications while preserving scene consistency. The edits are further integrated into the underlying 3D representation, supporting persistent and multi-view consistent novel-view rendering. Extensive experiments across both NeRF and 3D Gaussian Splatting backbones demonstrate that DesignAgent3D significantly outperforms state-of-the-art baselines, delivering superior semantic intent alignment, impeccable spatial localization accuracy, and high-fidelity multi-view consistency.
Problem

Research questions and friction points this paper is trying to address.

3D Scene Editing
Natural Language Design Requests
Spatial Grounding
Semantic Underspecification
Object Localization Drift
Innovation

Methods, ideas, or system contributions that make the work stand out.

Interactive Multimodal Reasoning
Plan-Perceive-Act Paradigm
Semantic Intent Alignment
Spatial Localization Accuracy
Multi-view Consistency
🔎 Similar Papers
No similar papers found.