SceneReGen: Generative Reconstruction of 3D Scenes from a Single Image

📅 2026-08-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
针对单图像3D场景重建中物体完成度与场景一致性问题,SceneReGen通过选择性姿态分解和预训练3D生成器生成完整物体并在共享场景框架中组装它们。
📝 Abstract
Single-image 3D scene reconstruction must complete partially observed objects and place them coherently in a shared observation-aligned scene frame. Object-level generative priors offer strong completion ability, but their centered, scale-normalized outputs are typically expressed in an object frame, creating a fundamental representation gap between object generation and scene reconstruction. We introduce SceneReGen, a generative reconstruction framework that reinterprets scene reconstruction as the generation and assembly of complete object assets in a shared observation-aligned scene frame. SceneReGen addresses the generation-reconstruction gap through selective pose factorization: each object's observed orientation is encoded directly in the generated mesh, while translation and scale are estimated from instance-level and global scene evidence. Given a scene image and instance masks, a geometry encoder extracts dense cues; learnable shape queries condition a pretrained DiT-based 3D generator to produce complete meshes in their observed orientations, while position queries fuse object and scene features to assemble them in the shared frame. On the 3D-FUTURE evaluation subset, SceneReGen achieves the best scene-level CD, scene-level F-Score, and 3D bounding-box IoU among the evaluated methods, ties the best object-level CD, and ranks second in object-level F-Score. Qualitative outputs in autonomous-driving and embodied-AI scenes further illustrate the potential of asset-centric reconstruction beyond indoor furniture.
Problem

Research questions and friction points this paper is trying to address.

Single-image 3D scene reconstruction
Object-level generative priors
Representation gap
Scene frame
Observation-aligned
Innovation

Methods, ideas, or system contributions that make the work stand out.

generative reconstruction
selective pose factorization
observation-aligned scene frame
shape queries
position queries
💼 Related Jobs
No related jobs found.
Z
Zefan Tian
Huawei
Y
Yuteng Ye
Huawei
Y
Yiheng Zhang
Northwestern Polytechnical University
Y
Yuhang Yang
Northwestern Polytechnical University
X
Xueqiang Lv
Northwestern Polytechnical University
Shizhou Zhang
Shizhou Zhang
Northwestern Polytechnical University
computer visionmachine learning
Le Liu
Le Liu
Northwestern Polytechnical University
VisualizationComputer GraphicsComputer VisionAI
Di Xu
Di Xu
Professor
Economics of EducationHigher Education PolicyProgram EvaluationCommunity CollegesOnline