RISE: Roadside Infrastructure Sequence Understanding across 3D Tracking and Structured Vision-Language Reasoning

📅 2026-08-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of LiDAR-free roadside 3D tracking and structured visual-language reasoning by proposing a framework integrating pure-vision tracking with constrained multimodal large language models. Cross-intersection tracking generalization is achieved through calibration-guided mask consistency, alongside the construction of a leakage-proof predictive VQA benchmark. Experimental results demonstrate that multi-view tracking achieves a MOTA of 66.9. Furthermore, the released RISE-VQA dataset, comprising 33,000 QA pairs, effectively reveals cognitive challenges in spatial localization and significantly advances roadside perception and reasoning capabilities.
📝 Abstract
We present RISE (Roadside Infrastructure Sequence Understanding and Evaluation), a framework spanning metric 3D tracking and structured vision-language reasoning in roadside sequences. For metric tracking, our image-only method combines SAM3 video identities with calibration-guided mask agreement for multi-view identity association, recovering persistent 3D tracks without LiDAR or task-specific 3D training. Its calibration-conditioned geometry allows the procedure to be instantiated at different calibrated multi-camera intersections without layout-specific retraining. On 20 human-reviewed clips from six intersections, the generated tracks achieve 66.9 MOTA within the defined multi-view evaluation scope. For structured vision-language reasoning, a human-reviewed MLLM pipeline mines high-value clips and uses a constrained full-context Oracle to construct bbox-grounded predictive QA without exposing future evidence to evaluated models. The resulting RISE-VQA dataset contains 33,910 QA pairs from 557 clips across 16 intersections and 61 roadside views. Its intersection-held-out RISE-Bench evaluates semantic choices, coordinates, future boxes, and interaction sets with deterministic task-specific metrics. Experiments show consistent benefits from domain adaptation and generally from temporal context, while revealing persistent challenges in spatial grounding, future localization, and interaction reasoning.
Problem

Research questions and friction points this paper is trying to address.

Roadside Infrastructure
3D Tracking
Vision-Language Reasoning
Benchmark Dataset
Spatial Grounding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Metric 3D Tracking
Structured Vision-Language Reasoning
Calibration-guided Mask Agreement
RISE-VQA
RISE-Bench