Federated Learning for Video Violence Detection: Complementary Roles of Lightweight CNNs and Vision-Language Models for Energy-Efficient Use

📅 2025-11-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
To address privacy preservation, high energy consumption, and non-IID data challenges in federated video violence detection, this paper proposes a privacy–energy-efficiency co-optimization framework. It integrates a lightweight 3D CNN with a vision-language model (LLaVA-NeXT-Video-7B), marking the first comparative study of LoRA-finetuned VLMs versus personalized CNNs in federated settings. A hierarchical semantic classification mechanism enhances multi-class recognition, while a scene-complexity-aware dynamic model invocation strategy enables zero-shot inference and semantic-aware class grouping. Experiments demonstrate a binary classification accuracy exceeding 90%, a 58% reduction in 3D CNN energy consumption (from 240 Wh to 570 Wh), an 81% multi-class accuracy for the VLM, and a ROC AUC of 92.59%. The framework thus achieves a balanced trade-off among detection accuracy, system sustainability, and deployment flexibility.

Technology Category

Application Category

📝 Abstract
Deep learning-based video surveillance increasingly demands privacy-preserving architectures with low computational and environmental overhead. Federated learning preserves privacy but deploying large vision-language models (VLMs) introduces major energy and sustainability challenges. We compare three strategies for federated violence detection under realistic non-IID splits on the RWF-2000 and RLVS datasets: zero-shot inference with pretrained VLMs, LoRA-based fine-tuning of LLaVA-NeXT-Video-7B, and personalized federated learning of a 65.8M-parameter 3D CNN. All methods exceed 90% accuracy in binary violence detection. The 3D CNN achieves superior calibration (ROC AUC 92.59%) at roughly half the energy cost (240 Wh vs. 570 Wh) of federated LoRA, while VLMs provide richer multimodal reasoning. Hierarchical category grouping (based on semantic similarity and class exclusion) boosts VLM multiclass accuracy from 65.31% to 81% on the UCF-Crime dataset. To our knowledge, this is the first comparative simulation study of LoRA-tuned VLMs and personalized CNNs for federated violence detection, with explicit energy and CO2e quantification. Our results inform hybrid deployment strategies that default to efficient CNNs for routine inference and selectively engage VLMs for complex contextual reasoning.
Problem

Research questions and friction points this paper is trying to address.

Comparing energy-efficient federated learning methods for video violence detection
Evaluating lightweight CNNs versus vision-language models under privacy constraints
Optimizing accuracy-energy tradeoffs in multimodal violence detection systems
Innovation

Methods, ideas, or system contributions that make the work stand out.

Federated learning with lightweight CNNs for energy efficiency
LoRA fine-tuning of vision-language models for multimodal reasoning
Hierarchical category grouping to boost VLM multiclass accuracy
💼 Related Jobs
No related jobs found.
S
Sébastien Thuau
esieaLab, ETIS Laboratory, ESIEA, University of CY Cergy, Paris, France
S
Siba Haidar
esieaLab, ESIEA, Paris, France
R
Rachid Chelouah
ETIS Laboratory, CNR1S, UMR8051, University of CY Cergy, Paris, France