🤖 AI Summary
This study addresses the conflation in existing literature among multimodal capabilities, agent performance, and deployment readiness, highlighting a lack of systematic evaluation of large models in real-world intelligent transportation systems (ITS). The work proposes the first large-scale, multimodal agent evidence assessment framework tailored for ITS, synthesizing insights from 42 research families published between January 2023 and August 2026. The framework distinguishes model-level, system-level, and hybrid architectures through functional tiers (C0–C3), validation levels (E0–E4), and eight methodological concern domains. Findings reveal that 23 studies support semantic processing and 24 achieve multidimensional fusion, yet only one reaches validation level E3. Large models demonstrate suitability for high-level interpretation and coordination tasks, but low-level control and safety-critical functions still require specialized systems or human oversight. A comparative evaluation protocol and a phased deployment roadmap are accordingly introduced.
📝 Abstract
Large multimodal agents (LMAs) are increasingly proposed for intelligent transportation systems (ITS), but existing studies often conflate multimodality, agency, empirical performance, and deployment readiness. This review provides an auditable evidence map of 42 primary study families released between January 2023 and 3 August 2026 within a corpus of 91 mapped sources. It distinguishes model-level, system-level, and hybrid multimodality and classifies each family by system architecture and action authority. Evidence is assessed independently through functional capability (C0-C3), validation setting (E0-E4), three evidence propositions (P1-P3), and eight methodological-concern domains (Q1-Q8). Transportation semantics (P1) are directly evaluated in 23 families and multidimensional integration (P3) in 24; 19 families directly evaluate both. Evidence reconciliation (P2) remains unresolved because no family demonstrates the complete provenance-challenge-handling-comparison-outcome chain. Fourteen families reach C3, but 13 remain at E2; only one reaches E3 and none reaches E4. Across ITS domains, LMAs are best supported for semantic interpretation, intent translation, evidence organisation, scenario authoring, explanation, and specialist-tool coordination. Numerical forecasting, optimisation, simulation fidelity, hard constraints, low-level control, safety fallback, and final authority should remain with independently verifiable specialist systems or accountable humans. The review therefore supports bounded orchestration rather than replacement and provides a matched comparative evaluation protocol and staged roadmap for accountable deployment. The living evidence repository is available at https://github.com/pangjunbiao/ITS-LMA-Review.