Generation of High-Level Concepts in 3D Scene Graphs via Autoregressive Diffusion

📅 2026-08-28
📈 Citations: 0
✹ Influential: 0
📄 PDF
🀖 AI Summary
本文提出䞀种基于自回園扩散的囟生成暡型甚于从观测到的垂盎平面自䞋而䞊构建3D场景囟解决了现有方法圚生成高级抂念时对特定规则和单独暡型䟝赖的问题。
📝 Abstract
Indoor 3D Scene Graphs (3DSGs) represent environments as multi-layer hierarchies that connect observed geometric primitives (e.g., planes) to higher-level metric-semantic concepts (e.g., rooms, floors, buildings), enabling incremental spatial reasoning for robotic perception and SLAM. However, classical high-level concept generation approaches rely on hand-crafted rules for specific concept classes, while learning-based methods require separate models for graph structure and spatial node features (e.g., centroids), which limits scalability to novel classes and more complex hierarchies. We propose a unified autoregressive diffusion-based graph generative model that jointly learns structure and features, constructing complete 3DSGs bottom-up from observed vertical planes across arbitrary hierarchy depths. Our method consistently surpasses all learning-based and random baselines across 3DSG datasets spanning synthetic scenes, real architectural floor plans, and robotic sensor data, with varying layout complexity and hierarchy depth, and surpasses a one-shot model with oracle access to the target graph size on the largest hierarchy and on real single-floor data. Finally, we propose an adaptation of the Fused Gromov--Wasserstein distance for principled graph-level evaluation of generated 3DSGs against ground truth.
Problem

Research questions and friction points this paper is trying to address.

3D Scene Graphs
High-Level Concepts
Autoregressive Diffusion
Graph Generative Model
Spatial Reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

autoregressive diffusion
unified generative model
3D Scene Graphs
bottom-up construction
Fused Gromov--Wasserstein distance
🔎 Similar Papers
2024-09-12arXiv.orgCitations: 9