Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Lucida通过解析、生成和放置方法解决真实场景到仿真环境的建模问题,提高3D对象检测与姿态估计精度。
📝 Abstract
Composable scene modeling aims to recover a real indoor scene as complete, editable object assets arranged as observed, giving robot simulation and embodied AI a simulation-ready replica of the real environment whose objects can be manipulated individually. Existing pipelines decompose the task into three steps---parse the observations into instances, generate an asset for each, and place each asset back---but every step presumes an input that a cluttered capture rarely provides: accurate instance geometry, unoccluded views, and assets that accurately match the observations. We propose Lucida, which keeps this order but redistributes the requirements, so each step consumes only what a real capture reliably provides and precision is reached at the end of the pipeline rather than demanded at its start. Lucida parses the video into a scene graph whose nodes carry per-instance multi-view evidence, generates a complete asset for each instance from its evidence, and places assets with GizmoAct, a VLM policy that casts placement as multi-turn GUI interaction, manipulating the object's gizmo in a closed loop and deciding itself when alignment is reached. Across scene-level 3D object detection, object pose estimation, and scene reconstruction, Lucida improves mAP over Boxer by 69% on R2S-Scene, raises ADD-SB@0.05 from 57.8% to 83.4% on CA-1M, and increases scene F-Score from 0.794 for SAM3D to 0.924.
Problem

Research questions and friction points this paper is trying to address.

composable scene modeling
instance geometry
unoccluded views
asset matching
Innovation

Methods, ideas, or system contributions that make the work stand out.

Composable Scene Modeling
Scene Graph
GizmoAct
VLM Policy
💼 Related Jobs
No related jobs found.
Minghan Qin
Minghan Qin
Bytedance Research | Tsinghua University
Computer Vision3D Vision3D Scene Perception
Yuang Wang
Yuang Wang
Zhejiang University
3D Vision3D Reconstruction
X
Xiuyu Yang
2Peking University
Y
Yushi Long
2Peking University
Y
Yujian Zhang
2Peking University
R
Ruihuan Wang
3Zhejiang University
K
Kai Ye
3Zhejiang University
Y
Yangang Zhang
3Zhejiang University
Hang Li
Hang Li
Bytedance Seed
natural language processinginformation retrievalmachine learningdata mining