Exploring 2D backbone effects for indoor semantic occupancy prediction

📅 2026-09-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过更换2D图像编码器(如DINOv2、BLIP2等)来提高室内语义占用预测的准确性,发现编码器的选择对最终3D预测有显著影响。
📝 Abstract
Semantic occupancy prediction gives an embodied agent a voxel-level account of where space is free, occupied, and semantically meaningful. In RGB-D pipelines such as EmbodiedScan, the image encoder is often left as a default module, even though its features are the visual evidence later sampled into the 3D grid. We study this design choice directly. A central finding is that changing the 2D backbone improves occupancy accuracy more than several carefully designed occupancy architectures or modules. We keep the main RGB-D projection, depth branch, and occupancy head fixed, and replace only the image backbone. The compared encoders are CLIP-ResNet, CLIP-ViT, BLIP2, and DINOv2. Under the controlled setting, the measured mIoU changes substantially: DINOv2 obtains 30.55\%, BLIP2 obtains 29.49\%, CLIP-ViT obtains 24.33\%, and CLIP-ResNet obtains 17.41\%. The stronger encoders also exceed the original EmbodiedScan ResNet-50 baseline without modifying the downstream 3D fusion pipeline. Class-level results give a more detailed picture: DINOv2 is stronger on many layout and structural categories, whereas BLIP2 remains close on several object-centered classes. CLIP-ViT improves clearly over CLIP-ResNet, showing that the way CLIP features are exposed as dense tokens matters for voxel lifting. These results indicate that the image backbone is not a secondary engineering detail in embodied semantic occupancy, but a major source of variation in the final 3D prediction.
Problem

Research questions and friction points this paper is trying to address.

Semantic Occupancy Prediction
2D Backbone
RGB-D Pipelines
Embodied Agent
Voxel-level
Innovation

Methods, ideas, or system contributions that make the work stand out.

2D Backbone
Semantic Occupancy Prediction
Image Encoder
Voxel-level Prediction
mIoU Improvement
💼 Related Jobs
No related jobs found.
S
Shizhang Fang
College of Electronics and Information Engineering, Shenzhen University, Shenzhen, China
W
Wanling Ye
College of Electronics and Information Engineering, Shenzhen University, Shenzhen, China
Qi Zheng
Qi Zheng
Shenzhen University
artificial intelligencemachine learning