BodyCam-VQA: Enhanced Body-Worn Camera Video Captioning via Multimodal Reasoning and Probe Question Generation
为解决警察随身摄像机视频数据处理难题,提出一种基于自适应视觉问答框架的方法,通过结构化推理和问题生成模型增强视频理解能力。
为解决警察随身摄像机视频数据处理难题,提出一种基于自适应视觉问答框架的方法,通过结构化推理和问题生成模型增强视频理解能力。
This study addresses the tight coupling between output paradigms and model architectures, datasets, and training protocols in existing video temporal localization methods, which hinders isolated evaluation of their impact on performance and lacks systematic analysis of deployment efficiency. Under unified experimental conditions—employing compact vision-language models (SmolVLM2, FastVLM, Molmo2) and LoRA fine-tuning—the work conducts controlled comparisons across three output paradigms: textual numeral generation, temporal token generation, and continuous temporal decoding on Charades-STA, QVHighlights, and YouCook2. Results demonstrate that continuous temporal decoding consistently achieves the best accuracy-efficiency trade-off along the Pareto frontier, maintaining high localization accuracy while significantly reducing inference latency and parameter overhead, thereby offering clear design guidance for edge deployment.
This study addresses the challenge of generating trustworthy incident reports from multi-role, high-noise spoken dialogues in law enforcement settings. We propose the first trust-centered large language model (LLM) framework for this task. Methodologically, it integrates role-aware information extraction with procedural narrative generation, employing dialogue role separation, robust key-event extraction, noise-aware prompt engineering, and structured output constraints to safeguard officer and civilian rights while ensuring regulatory compliance. Our key contribution lies in explicitly modeling accountability, fairness, and transparency as primary LLM generation objectives—rather than as post-hoc attributes. Evaluated on real-world police ASR transcripts, the framework achieves 92% recall for critical report elements and 87% procedural correctness, significantly improving report consistency and auditability. This work establishes a verifiable, deployable technical paradigm for intelligent policing documentation systems.
为解决警察随身摄像机视频数据处理难题,提出一种基于自适应视觉问答框架的方法,通过结构化推理和问题生成模型增强视频理解能力。
This study addresses the tight coupling between output paradigms and model architectures, datasets, and training protocols in existing video temporal localization methods, which hinders isolated evaluation of their impact on performance and lacks systematic analysis of deployment efficiency. Under unified experimental conditions—employing compact vision-language models (SmolVLM2, FastVLM, Molmo2) and LoRA fine-tuning—the work conducts controlled comparisons across three output paradigms: textual numeral generation, temporal token generation, and continuous temporal decoding on Charades-STA, QVHighlights, and YouCook2. Results demonstrate that continuous temporal decoding consistently achieves the best accuracy-efficiency trade-off along the Pareto frontier, maintaining high localization accuracy while significantly reducing inference latency and parameter overhead, thereby offering clear design guidance for edge deployment.
This study addresses the challenge of generating trustworthy incident reports from multi-role, high-noise spoken dialogues in law enforcement settings. We propose the first trust-centered large language model (LLM) framework for this task. Methodologically, it integrates role-aware information extraction with procedural narrative generation, employing dialogue role separation, robust key-event extraction, noise-aware prompt engineering, and structured output constraints to safeguard officer and civilian rights while ensuring regulatory compliance. Our key contribution lies in explicitly modeling accountability, fairness, and transparency as primary LLM generation objectives—rather than as post-hoc attributes. Evaluated on real-world police ASR transcripts, the framework achieves 92% recall for critical report elements and 87% procedural correctness, significantly improving report consistency and auditability. This work establishes a verifiable, deployable technical paradigm for intelligent policing documentation systems.