UniTraffic-Agent: Unified Traffic Video Reasoning for AI City Challenge 2026 Track 3 with Two Out-of-Domain Evaluations

📅 2026-08-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenges of traffic video understanding—such as sparse events and highly variable viewpoints—that hinder accurate parsing of event evolution and critical interactions. We propose UniTraffic-Agent, the first unified agent framework for multi-task traffic video reasoning, integrating traffic anomaly detection, fisheye video event understanding, and pedestrian intent question answering. Built upon a multimodal large language model, our approach employs an “observe–reason–act–verify” workflow, enhanced with timestamp-aware visual evidence sampling and task-specific action adapters to enable joint reasoning and cross-domain generalization. Evaluated in the AI City Challenge 2026, UniTraffic-Agent achieves leading performance across all three tasks: ranking 16th in TAR (0.5780), 2nd in FETV (0.4884), and 4th in PSI-VQA (64.4161).
📝 Abstract
Traffic video understanding has become an important problem in intelligent transportation, as road videos provide direct evidence for accidents, violations, and interactions between vehicles and vulnerable road users. A useful system should explain how a traffic event develops, why it happens, and when the relevant interaction occurs, yet this remains difficult for multimodal large language models (MLLMs) because traffic videos contain sparse events and varied viewpoints. We introduce UniTraffic-Agent, the MR-CAS solution for Track~3 of the 10th AI City Challenge, which includes Traffic Anomaly Reasoning (TAR) and two out-of-domain evaluations: FETV for fisheye traffic events and PSI-VQA for pedestrian intention reasoning. UniTraffic-Agent follows an observe--reason--act--verify workflow that samples timestamped visual evidence, reasons over all questions from the same clip in one request, and converts responses through task-specific action adapters. On the official Public leaderboards, MR-CAS ranks 16th on TAR with a score of 0.5780, 2nd on FETV with 0.4884, and 4th on PSI-VQA with 64.4161. The code is available at https://github.com/Roclp/UniTraffic-Agent.
Problem

Research questions and friction points this paper is trying to address.

traffic video understanding
multimodal large language models
traffic anomaly reasoning
out-of-domain evaluation
intelligent transportation
Innovation

Methods, ideas, or system contributions that make the work stand out.

UniTraffic-Agent
traffic video reasoning
multimodal large language models
out-of-domain generalization
action adapters
P
Peng Li
School of Computer Science and Technology, University of Chinese Academy of Sciences (UCAS), Beijing, China
Q
Qianqian Xu
State Key Laboratory of AI Safety, Institute of Computing Technology (ICT), Chinese Academy of Sciences (CAS), Beijing, China; Beijing Academy of Artificial Intelligence (BAAI), Beijing, China
S
Shilong Bao
State Key Laboratory of AI Safety, Institute of Computing Technology (ICT), Chinese Academy of Sciences (CAS), Beijing, China
Yangbangyan Jiang
Yangbangyan Jiang
University of Chinese Academy of Sciences
Machine LearningDeep learning
Qingming Huang
Qingming Huang
University of the Chinese Academy of Sciences
Multimedia Analysis and RetrievalImage and Video ProcessingPattern RecognitionComputer VisionVideo Coding