Thinking with Cameras: Active Visual Reasoning via Dynamic Viewpoint Control for Surveillance Video Understanding
为解决监控视频理解中目标远、小、遮挡或移出视野的问题,本文通过动态视角控制和强化学习优化视角策略的方法,提高视觉证据获取能力。
为解决监控视频理解中目标远、小、遮挡或移出视野的问题,本文通过动态视角控制和强化学习优化视角策略的方法,提高视觉证据获取能力。
为解决跨平台3D目标检测中因传感器高度和视角变化导致的点分布改变问题,提出SimFuse3D方法,通过源指导目标模拟与置信度引导的多阶段定位重加权来改善预测准确性。
研究统一多模态模型中理解和生成任务的协同效应,通过调整架构和任务知识共享,促进两者从共存到协同。
本文通过引入视觉渲染作为上下文管理器,解决了长时序代理在多模态历史下的上下文压缩问题,并提出VERA策略以保留原生视觉证据。
This work addresses the limitations of existing industrial safety datasets, which are typically confined to single-modality perception or isolated violation detection and thus incapable of supporting evidence-based, multi-step reasoning for compliance assessment, accident mechanism analysis, and preventive recommendations. To bridge this gap, we introduce SafeSceneReason—the first multimodal industrial safety reasoning benchmark that integrates accident investigation knowledge. Our approach employs a dual-track pipeline centered on scenes and reports to align workplace images with accident narratives, generating question-answer pairs spanning perception, compliance judgment, causal analysis, and actionable recommendations. The benchmark innovatively combines executable safety scene graphs with accident evidence graphs, leveraging procedural execution, evidence extraction, and multi-hop reasoning path generation to construct a high-quality dataset of 123,695 question-answer pairs. Evaluations reveal that current vision-language models exhibit significant deficiencies in technical, comparative, and multi-evidence reasoning tasks.
为解决监控视频理解中目标远、小、遮挡或移出视野的问题,本文通过动态视角控制和强化学习优化视角策略的方法,提高视觉证据获取能力。
为解决跨平台3D目标检测中因传感器高度和视角变化导致的点分布改变问题,提出SimFuse3D方法,通过源指导目标模拟与置信度引导的多阶段定位重加权来改善预测准确性。
研究统一多模态模型中理解和生成任务的协同效应,通过调整架构和任务知识共享,促进两者从共存到协同。
本文通过引入视觉渲染作为上下文管理器,解决了长时序代理在多模态历史下的上下文压缩问题,并提出VERA策略以保留原生视觉证据。
This work addresses the limitations of existing industrial safety datasets, which are typically confined to single-modality perception or isolated violation detection and thus incapable of supporting evidence-based, multi-step reasoning for compliance assessment, accident mechanism analysis, and preventive recommendations. To bridge this gap, we introduce SafeSceneReason—the first multimodal industrial safety reasoning benchmark that integrates accident investigation knowledge. Our approach employs a dual-track pipeline centered on scenes and reports to align workplace images with accident narratives, generating question-answer pairs spanning perception, compliance judgment, causal analysis, and actionable recommendations. The benchmark innovatively combines executable safety scene graphs with accident evidence graphs, leveraging procedural execution, evidence extraction, and multi-hop reasoning path generation to construct a high-quality dataset of 123,695 question-answer pairs. Evaluations reveal that current vision-language models exhibit significant deficiencies in technical, comparative, and multi-evidence reasoning tasks.