Institution profile

Milestone Systems

Industry researcheurope · dk
Official website
Research library10linked papers
Opportunities0open roles
Selected work

Representative Papers

The 10th AI City Challenge

Aug 17, 2026

本文总结了第10届AI City挑战赛,该赛事通过结合基础模型、几何定位等方法解决智能交通和智慧城市中的多摄像头感知等问题。

0 citationsRead paper

MVTrack: Ultrafast Appearance-Free Moving Object Tracking from Compressed Bitstreams

Aug 11, 2026

This work addresses the high computational cost of existing RGB-based video object tracking methods, which hinders their large-scale deployment. The authors propose MVTrack, the first approach to achieve efficient tracking solely using motion vectors extracted from H.264 compressed bitstreams, entirely bypassing pixel-domain processing. MVTrack integrates a lightweight motion vector field detector (MVDet) with a minimalistic motion association module (MVLink) to enable accurate tracking without video decoding. Evaluated on the VIRAT dataset, MVTrack outperforms YOLOv2tiny in tracking accuracy while using 60× fewer parameters, requiring 40× lower FLOPs, and achieving 8.6× faster CPU inference speed, thereby significantly advancing the practicality and efficiency of compressed-domain tracking.

0 citationsRead paper

From Pixels to Privacy: Temporally Consistent Video Anonymization via Token Pruning for Privacy Preserving Action Recognition

Mar 27, 2026

This work addresses the privacy risks inherent in video understanding models, which, while achieving high action recognition performance, often inadvertently leak sensitive attributes such as identity, gender, and race. To mitigate this, the authors propose a spatiotemporal anonymization framework based on Vision Transformers that introduces dual classification tokens—dedicated to action and privacy—within a unified architecture. By contrasting the attention distributions of these tokens, the method establishes a utility-privacy scoring mechanism to identify and prune low-scoring spatiotemporal tubelets. This approach effectively disentangles utility-related and privacy-sensitive features, leveraging attention discrepancies to guide anonymization. Extensive experiments demonstrate that the framework preserves action recognition accuracy comparable to that of the original videos while significantly reducing the leakage of sensitive attributes across multiple benchmarks.

0 citationsRead paper

Only Whats Necessary: Pareto Optimal Data Minimization for Privacy Preserving Video Anomaly Detection

Mar 27, 2026

This work addresses the challenge of complying with GDPR’s data minimization principle in video anomaly detection, where personal identifiable information often complicates privacy compliance. To this end, the authors propose a privacy-first design framework that, for the first time, introduces Pareto optimality into this domain. The approach integrates a breadth-and-depth data minimization mechanism that effectively suppresses sensitive visual content while preserving essential cues for anomaly detection. It employs visual information suppression techniques, a dual-model evaluation architecture comprising an anomaly detection model and a privacy inference model, and a rank-based assessment strategy. Pareto front analysis is used to quantitatively characterize the trade-off between privacy preservation and task utility, enabling the identification of optimal operating points. Experiments on public datasets demonstrate that the framework substantially reduces exposure of personal data with only marginal performance degradation.

0 citationsRead paper

SOVABench: A Vehicle Surveillance Action Retrieval Benchmark for Multimodal Large Language Models

Jan 08, 2026arXiv.org

Existing video retrieval benchmarks primarily focus on scene-level similarity, which is insufficient for evaluating fine-grained discriminative capabilities in surveillance scenarios involving vehicle actions. To address this limitation, this work proposes SOVABench, the first vehicle behavior retrieval benchmark tailored to real-world surveillance settings, and introduces two novel evaluation protocols to assess models’ understanding of action oppositionality and temporal directionality. Leveraging the visual reasoning and instruction-following capabilities of multimodal large language models (MLLMs), the proposed method generates interpretable textual embeddings for zero-shot image-to-video retrieval without requiring task-specific training. Experimental results demonstrate that this approach significantly outperforms conventional contrastive vision-language models on SOVABench as well as multiple spatial and counting benchmarks, confirming its effectiveness and strong generalization ability.

0 citationsRead paper
Recent publications

Latest Papers

The 10th AI City Challenge

Aug 17, 2026

本文总结了第10届AI City挑战赛,该赛事通过结合基础模型、几何定位等方法解决智能交通和智慧城市中的多摄像头感知等问题。

0 citationsRead paper

MVTrack: Ultrafast Appearance-Free Moving Object Tracking from Compressed Bitstreams

Aug 11, 2026

This work addresses the high computational cost of existing RGB-based video object tracking methods, which hinders their large-scale deployment. The authors propose MVTrack, the first approach to achieve efficient tracking solely using motion vectors extracted from H.264 compressed bitstreams, entirely bypassing pixel-domain processing. MVTrack integrates a lightweight motion vector field detector (MVDet) with a minimalistic motion association module (MVLink) to enable accurate tracking without video decoding. Evaluated on the VIRAT dataset, MVTrack outperforms YOLOv2tiny in tracking accuracy while using 60× fewer parameters, requiring 40× lower FLOPs, and achieving 8.6× faster CPU inference speed, thereby significantly advancing the practicality and efficiency of compressed-domain tracking.

0 citationsRead paper

From Pixels to Privacy: Temporally Consistent Video Anonymization via Token Pruning for Privacy Preserving Action Recognition

Mar 27, 2026

This work addresses the privacy risks inherent in video understanding models, which, while achieving high action recognition performance, often inadvertently leak sensitive attributes such as identity, gender, and race. To mitigate this, the authors propose a spatiotemporal anonymization framework based on Vision Transformers that introduces dual classification tokens—dedicated to action and privacy—within a unified architecture. By contrasting the attention distributions of these tokens, the method establishes a utility-privacy scoring mechanism to identify and prune low-scoring spatiotemporal tubelets. This approach effectively disentangles utility-related and privacy-sensitive features, leveraging attention discrepancies to guide anonymization. Extensive experiments demonstrate that the framework preserves action recognition accuracy comparable to that of the original videos while significantly reducing the leakage of sensitive attributes across multiple benchmarks.

0 citationsRead paper

Only Whats Necessary: Pareto Optimal Data Minimization for Privacy Preserving Video Anomaly Detection

Mar 27, 2026

This work addresses the challenge of complying with GDPR’s data minimization principle in video anomaly detection, where personal identifiable information often complicates privacy compliance. To this end, the authors propose a privacy-first design framework that, for the first time, introduces Pareto optimality into this domain. The approach integrates a breadth-and-depth data minimization mechanism that effectively suppresses sensitive visual content while preserving essential cues for anomaly detection. It employs visual information suppression techniques, a dual-model evaluation architecture comprising an anomaly detection model and a privacy inference model, and a rank-based assessment strategy. Pareto front analysis is used to quantitatively characterize the trade-off between privacy preservation and task utility, enabling the identification of optimal operating points. Experiments on public datasets demonstrate that the framework substantially reduces exposure of personal data with only marginal performance degradation.

0 citationsRead paper

SOVABench: A Vehicle Surveillance Action Retrieval Benchmark for Multimodal Large Language Models

Jan 08, 2026arXiv.org

Existing video retrieval benchmarks primarily focus on scene-level similarity, which is insufficient for evaluating fine-grained discriminative capabilities in surveillance scenarios involving vehicle actions. To address this limitation, this work proposes SOVABench, the first vehicle behavior retrieval benchmark tailored to real-world surveillance settings, and introduces two novel evaluation protocols to assess models’ understanding of action oppositionality and temporal directionality. Leveraging the visual reasoning and instruction-following capabilities of multimodal large language models (MLLMs), the proposed method generates interpretable textual embeddings for zero-shot image-to-video retrieval without requiring task-specific training. Experimental results demonstrate that this approach significantly outperforms conventional contrastive vision-language models on SOVABench as well as multiple spatial and counting benchmarks, confirming its effectiveness and strong generalization ability.

0 citationsRead paper