Institution profile

SenseTime

Industry researchasia · cn
Official website
Research library60linked papers
Opportunities0open roles
Selected work

Representative Papers

SafeSceneReason: A Multimodal Reasoning Benchmark Connecting Industrial Hazards with Accident Knowledge

Aug 10, 2026

This work addresses the limitations of existing industrial safety datasets, which are typically confined to single-modality perception or isolated violation detection and thus incapable of supporting evidence-based, multi-step reasoning for compliance assessment, accident mechanism analysis, and preventive recommendations. To bridge this gap, we introduce SafeSceneReason—the first multimodal industrial safety reasoning benchmark that integrates accident investigation knowledge. Our approach employs a dual-track pipeline centered on scenes and reports to align workplace images with accident narratives, generating question-answer pairs spanning perception, compliance judgment, causal analysis, and actionable recommendations. The benchmark innovatively combines executable safety scene graphs with accident evidence graphs, leveraging procedural execution, evidence extraction, and multi-hop reasoning path generation to construct a high-quality dataset of 123,695 question-answer pairs. Evaluations reveal that current vision-language models exhibit significant deficiencies in technical, comparative, and multi-evidence reasoning tasks.

0 citationsRead paper
Recent publications

Latest Papers

SafeSceneReason: A Multimodal Reasoning Benchmark Connecting Industrial Hazards with Accident Knowledge

Aug 10, 2026

This work addresses the limitations of existing industrial safety datasets, which are typically confined to single-modality perception or isolated violation detection and thus incapable of supporting evidence-based, multi-step reasoning for compliance assessment, accident mechanism analysis, and preventive recommendations. To bridge this gap, we introduce SafeSceneReason—the first multimodal industrial safety reasoning benchmark that integrates accident investigation knowledge. Our approach employs a dual-track pipeline centered on scenes and reports to align workplace images with accident narratives, generating question-answer pairs spanning perception, compliance judgment, causal analysis, and actionable recommendations. The benchmark innovatively combines executable safety scene graphs with accident evidence graphs, leveraging procedural execution, evidence extraction, and multi-hop reasoning path generation to construct a high-quality dataset of 123,695 question-answer pairs. Evaluations reveal that current vision-language models exhibit significant deficiencies in technical, comparative, and multi-evidence reasoning tasks.

0 citationsRead paper