Inter-3D VQA: A Roadside Multimodal Benchmark for 3D Spatiotemporally Grounded Visual Question Answering

📅 2026-08-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出Inter-3D VQA,一个基于多模态的3D时空基准,用于解决交通场景中的视觉问答问题,并通过整合LiDAR表示法来改进模型性能。
📝 Abstract
Recent advances in visual question answering (VQA) and multimodal large language models (MLLMs) have enabled natural-language reasoning over traffic scenes. However, existing benchmarks are largely built from ego-vehicle views or 2D roadside videos, limiting their ability to evaluate 3D-grounded reasoning over real-world distances, trajectories, infrastructure topology, and safety-critical interactions. We introduce Inter-3D VQA, a large-scale roadside multimodal benchmark for 3D spatiotemporally grounded VQA at intersections. Built from synchronized point clouds and multi-view images, Inter-3D VQA contains 407K QA pairs covering lane-level positions, object relationships, motion patterns, and near-miss-oriented interaction reasoning. We further propose Inter-Geo, an MLLM baseline that integrates object- and scene-level aligned LiDAR representations, and Inter-Metrics, a unified evaluation framework for textual consistency, numerical accuracy, and semantic correctness. Experiments show that Inter-Geo outperforms image-based VLMs, especially on grounded spatial and temporal reasoning tasks. Our benchmark and codes are available at https://github.com/ASU-Suo-Lab/Inter-3D-VQA .
Problem

Research questions and friction points this paper is trying to address.

3D spatiotemporally grounded
visual question answering
traffic scenes
benchmarks
ego-vehicle views
Innovation

Methods, ideas, or system contributions that make the work stand out.

3D spatiotemporally grounded VQA
multimodal benchmark
LiDAR representations
unified evaluation framework
💼 Related Jobs
No related jobs found.
S
Shaozu Ding
The Polytechnic School, Arizona State University
L
Linan Song
The Polytechnic School, Arizona State University
Dajiang Suo
Dajiang Suo
Assistant Professor, Arizona State University
IoTMutimodal sensingCV cybersecurityIntelligent infrastructure