BodyCam-VQA: Enhanced Body-Worn Camera Video Captioning via Multimodal Reasoning and Probe Question Generation

๐Ÿ“… 2026-09-09
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
ไธบ่งฃๅ†ณ่ญฆๅฏŸ้š่บซๆ‘„ๅƒๆœบ่ง†้ข‘ๆ•ฐๆฎๅค„็†้šพ้ข˜๏ผŒๆๅ‡บไธ€็งๅŸบไบŽ่‡ช้€‚ๅบ”่ง†่ง‰้—ฎ็ญ”ๆก†ๆžถ็š„ๆ–นๆณ•๏ผŒ้€š่ฟ‡็ป“ๆž„ๅŒ–ๆŽจ็†ๅ’Œ้—ฎ้ข˜็”Ÿๆˆๆจกๅž‹ๅขžๅผบ่ง†้ข‘็†่งฃ่ƒฝๅŠ›ใ€‚
๐Ÿ“ Abstract
Police body-worn camera (BWC) footage has emerged as a critical aspect of law enforcement that ensures legal transparency, officer accountability, and the protection of civil rights. However, effectively processing this data remains a significant challenge due to its multimodal video format. BWC videos, in many cases, comprise chaotic scenes with low visual quality, rapid movement/interactions, and high-noise audio that make visual understanding a challenge for even SOTA multimodal models. Current Vision-Language Models (VLMs) frequently overlook critical forensic details, such as the presence of valuable evidence or the latent nuances of suspect-officer interactions, which are vital for fair legal outcomes and civilian/officer safety. To address these limitations, we propose an Adaptive Visual Question Answering (VQA) framework engineered for high-stakes law enforcement. Our framework employs a structured reasoning approach to extract fine-grained visual evidence that traditional captioning systems fail to capture. We experiment with multiple question generation models, including foundation models and fine-tuned open-weight models, to observe performance variation among question generation model implementations. Our results demonstrate that this VQA-driven architecture provides a more reliable, objective, and detailed record of enforcement events, ultimately serving as a powerful tool to protect both law enforcement officers and the public through AI-assisted forensic clarity.
Problem

Research questions and friction points this paper is trying to address.

Body-Worn Camera
Multimodal Video
Forensic Details
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adaptive Visual Question Answering (VQA)
structured reasoning
fine-grained visual evidence
police body-worn camera (BWC) footage
multimodal models
๐Ÿ”Ž Similar Papers
2024-02-20International Conference on Machine LearningCitations: 30
๐Ÿ’ผ Related Jobs
No related jobs found.
K
Karish Gupta
Worcester Polytechnic Institute, Worcester, MA, USA
M
Matthew Alex
Worcester Polytechnic Institute, Worcester, MA, USA
A
Alex Li
Worcester Polytechnic Institute, Worcester, MA, USA
Y
Yang Wu
Worcester Polytechnic Institute, Worcester, MA, USA
Yun-Wei Chu
Yun-Wei Chu
Purdue University
Machine LearningNatural Language Processing
K
Kashif Munir
Axon, Scottsdale, AZ, USA
Xiaotian Zhou
Xiaotian Zhou
Fudan University
graph data miningsocial networks
Z
Zhengping Ji
Axon, Scottsdale, AZ, USA
Xiaozhong Liu
Xiaozhong Liu
School of Informatics and Computing, Indiana University Bloomington
Information RetrievalNatural Language ProcessingDigital LibrarySemantic WebMetadata