The 10th AI City Challenge
本文总结了第10届AI City挑战赛,该赛事通过结合基础模型、几何定位等方法解决智能交通和智慧城市中的多摄像头感知等问题。
本文总结了第10届AI City挑战赛,该赛事通过结合基础模型、几何定位等方法解决智能交通和智慧城市中的多摄像头感知等问题。
This work addresses the high computational cost of existing RGB-based video object tracking methods, which hinders their large-scale deployment. The authors propose MVTrack, the first approach to achieve efficient tracking solely using motion vectors extracted from H.264 compressed bitstreams, entirely bypassing pixel-domain processing. MVTrack integrates a lightweight motion vector field detector (MVDet) with a minimalistic motion association module (MVLink) to enable accurate tracking without video decoding. Evaluated on the VIRAT dataset, MVTrack outperforms YOLOv2tiny in tracking accuracy while using 60× fewer parameters, requiring 40× lower FLOPs, and achieving 8.6× faster CPU inference speed, thereby significantly advancing the practicality and efficiency of compressed-domain tracking.
This work addresses the privacy risks inherent in video understanding models, which, while achieving high action recognition performance, often inadvertently leak sensitive attributes such as identity, gender, and race. To mitigate this, the authors propose a spatiotemporal anonymization framework based on Vision Transformers that introduces dual classification tokens—dedicated to action and privacy—within a unified architecture. By contrasting the attention distributions of these tokens, the method establishes a utility-privacy scoring mechanism to identify and prune low-scoring spatiotemporal tubelets. This approach effectively disentangles utility-related and privacy-sensitive features, leveraging attention discrepancies to guide anonymization. Extensive experiments demonstrate that the framework preserves action recognition accuracy comparable to that of the original videos while significantly reducing the leakage of sensitive attributes across multiple benchmarks.
This work addresses the challenge of complying with GDPR’s data minimization principle in video anomaly detection, where personal identifiable information often complicates privacy compliance. To this end, the authors propose a privacy-first design framework that, for the first time, introduces Pareto optimality into this domain. The approach integrates a breadth-and-depth data minimization mechanism that effectively suppresses sensitive visual content while preserving essential cues for anomaly detection. It employs visual information suppression techniques, a dual-model evaluation architecture comprising an anomaly detection model and a privacy inference model, and a rank-based assessment strategy. Pareto front analysis is used to quantitatively characterize the trade-off between privacy preservation and task utility, enabling the identification of optimal operating points. Experiments on public datasets demonstrate that the framework substantially reduces exposure of personal data with only marginal performance degradation.
Existing video retrieval benchmarks primarily focus on scene-level similarity, which is insufficient for evaluating fine-grained discriminative capabilities in surveillance scenarios involving vehicle actions. To address this limitation, this work proposes SOVABench, the first vehicle behavior retrieval benchmark tailored to real-world surveillance settings, and introduces two novel evaluation protocols to assess models’ understanding of action oppositionality and temporal directionality. Leveraging the visual reasoning and instruction-following capabilities of multimodal large language models (MLLMs), the proposed method generates interpretable textual embeddings for zero-shot image-to-video retrieval without requiring task-specific training. Experimental results demonstrate that this approach significantly outperforms conventional contrastive vision-language models on SOVABench as well as multiple spatial and counting benchmarks, confirming its effectiveness and strong generalization ability.
本文总结了第10届AI City挑战赛,该赛事通过结合基础模型、几何定位等方法解决智能交通和智慧城市中的多摄像头感知等问题。
This work addresses the high computational cost of existing RGB-based video object tracking methods, which hinders their large-scale deployment. The authors propose MVTrack, the first approach to achieve efficient tracking solely using motion vectors extracted from H.264 compressed bitstreams, entirely bypassing pixel-domain processing. MVTrack integrates a lightweight motion vector field detector (MVDet) with a minimalistic motion association module (MVLink) to enable accurate tracking without video decoding. Evaluated on the VIRAT dataset, MVTrack outperforms YOLOv2tiny in tracking accuracy while using 60× fewer parameters, requiring 40× lower FLOPs, and achieving 8.6× faster CPU inference speed, thereby significantly advancing the practicality and efficiency of compressed-domain tracking.
This work addresses the privacy risks inherent in video understanding models, which, while achieving high action recognition performance, often inadvertently leak sensitive attributes such as identity, gender, and race. To mitigate this, the authors propose a spatiotemporal anonymization framework based on Vision Transformers that introduces dual classification tokens—dedicated to action and privacy—within a unified architecture. By contrasting the attention distributions of these tokens, the method establishes a utility-privacy scoring mechanism to identify and prune low-scoring spatiotemporal tubelets. This approach effectively disentangles utility-related and privacy-sensitive features, leveraging attention discrepancies to guide anonymization. Extensive experiments demonstrate that the framework preserves action recognition accuracy comparable to that of the original videos while significantly reducing the leakage of sensitive attributes across multiple benchmarks.
This work addresses the challenge of complying with GDPR’s data minimization principle in video anomaly detection, where personal identifiable information often complicates privacy compliance. To this end, the authors propose a privacy-first design framework that, for the first time, introduces Pareto optimality into this domain. The approach integrates a breadth-and-depth data minimization mechanism that effectively suppresses sensitive visual content while preserving essential cues for anomaly detection. It employs visual information suppression techniques, a dual-model evaluation architecture comprising an anomaly detection model and a privacy inference model, and a rank-based assessment strategy. Pareto front analysis is used to quantitatively characterize the trade-off between privacy preservation and task utility, enabling the identification of optimal operating points. Experiments on public datasets demonstrate that the framework substantially reduces exposure of personal data with only marginal performance degradation.
Existing video retrieval benchmarks primarily focus on scene-level similarity, which is insufficient for evaluating fine-grained discriminative capabilities in surveillance scenarios involving vehicle actions. To address this limitation, this work proposes SOVABench, the first vehicle behavior retrieval benchmark tailored to real-world surveillance settings, and introduces two novel evaluation protocols to assess models’ understanding of action oppositionality and temporal directionality. Leveraging the visual reasoning and instruction-following capabilities of multimodal large language models (MLLMs), the proposed method generates interpretable textual embeddings for zero-shot image-to-video retrieval without requiring task-specific training. Experimental results demonstrate that this approach significantly outperforms conventional contrastive vision-language models on SOVABench as well as multiple spatial and counting benchmarks, confirming its effectiveness and strong generalization ability.