SceneBench: A Hierarchical Benchmark for Vision-Language Understanding of 3D Scenes

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决现有3D场景理解基准的局限性,本文通过构建包含966个光真实3D场景的SceneBench并定义三个评估任务来提升视觉-语言模型在3D空间推理方面的能力。
📝 Abstract
Vision-language models excel at 2D image understanding but remain limited in 3D spatial reasoning. Progress is hindered by limitations in current benchmarks. First, 3D datasets often rely on point clouds that capture geometry but discard rich visual features like texture, text, and materials. Second, annotations treat objects in isolation while ignoring real-world hierarchical organization (scenes, rooms, functional areas, object groups). Third, evaluation tasks focus narrowly on basic recognition rather than multi-step spatial reasoning. In this context, we introduce SceneBench, a benchmark of 966 photorealistic 3D scenes reconstructed with Gaussian Splatting and densely annotated with hierarchical semantics spanning scenes, rooms, functional areas, object groups, and individual objects. These annotations are produced through a human-in-the-loop pipeline combining vision-language models with roughly 1,500 human-hours of iterative refinement and verification, producing over 183K annotated nodes with textual descriptions and 3D bounding boxes. Building on this representation, we define three evaluation tasks: Existence-Based Questions probing object attributes, Spatial Intelligence Questions covering counting, size comparison, distance, and directional relations, and Grounded Question-Reasoning-Answer (QRA) triplets requiring multi-step reasoning across semantic levels. Experiments with state-of-the-art vision-language models show that while models perform well on basic recognition tasks (e.g., up to 85% accuracy for detection), performance drops substantially on hierarchical and compositional reasoning (e.g., down to 60% for counting), revealing limitations not captured by existing benchmarks. SceneBench provides a realistic testbed for developing and evaluating models capable of fine-grained spatial reasoning in photorealistic 3D environments.
Problem

Research questions and friction points this paper is trying to address.

3D spatial reasoning
hierarchical organization
multi-step reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hierarchical Semantics
Photorealistic 3D Scenes
Spatial Reasoning
Human-in-the-Loop Annotation
Multi-step Reasoning
A
Anubhav Khanal
Nepal Applied Mathematics and Informatics Institute for research (NAAMII), Nepal
P
Prabigya Acharya
Nepal Applied Mathematics and Informatics Institute for research (NAAMII), Nepal
R
Roshni Poudel
Nepal Applied Mathematics and Informatics Institute for research (NAAMII), Nepal
S
Sujan Kapali
Nepal Applied Mathematics and Informatics Institute for research (NAAMII), Nepal
B
Bigyan Bhatta
Nepal Applied Mathematics and Informatics Institute for research (NAAMII), Nepal
P
Pramish Paudel
INSAIT, Sofia University, Bulgaria
F
Francois Rameau
Nepal Applied Mathematics and Informatics Institute for research (NAAMII), Nepal; State University of New York at Stony Brook, Korea
Danda Pani Paudel
Danda Pani Paudel
INSAIT Sofia University
Computer VisionRoboticsEarth Observation