SafeSceneReason: A Multimodal Reasoning Benchmark Connecting Industrial Hazards with Accident Knowledge

📅 2026-08-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of existing industrial safety datasets, which are typically confined to single-modality perception or isolated violation detection and thus incapable of supporting evidence-based, multi-step reasoning for compliance assessment, accident mechanism analysis, and preventive recommendations. To bridge this gap, we introduce SafeSceneReason—the first multimodal industrial safety reasoning benchmark that integrates accident investigation knowledge. Our approach employs a dual-track pipeline centered on scenes and reports to align workplace images with accident narratives, generating question-answer pairs spanning perception, compliance judgment, causal analysis, and actionable recommendations. The benchmark innovatively combines executable safety scene graphs with accident evidence graphs, leveraging procedural execution, evidence extraction, and multi-hop reasoning path generation to construct a high-quality dataset of 123,695 question-answer pairs. Evaluations reveal that current vision-language models exhibit significant deficiencies in technical, comparative, and multi-evidence reasoning tasks.
📝 Abstract
Industrial-safety understanding requires more than detecting workers, equipment, and personal protective equipment. Models must also assess compliance, identify hazardous interactions, explain potential accident mechanisms, and recommend preventive actions. Existing safety datasets primarily focus on visual perception or isolated violation recognition and provide limited supervision for evidence-grounded reasoning. We introduce SafeSceneReason, a multimodal industrial-safety reasoning benchmark and companion training corpus that connects workplace scenes with knowledge from occupational accident investigations. SafeSceneReason combines two complementary data-construction pipelines. The scene-centric pipeline converts annotated workplace images into executable safety scene graphs and generates deterministic answers through program execution over objects, relations, and safety rules. The report-centric pipeline extracts figures and contextual evidence from accident reports and constructs multimodal questions using evidence graphs, explicit information boundaries, multi-step reasoning paths, and iterative verification. The resulting resource contains 110,581 verified scene-centric question--answer pairs and 13,114 refined report-centric question--answer pairs, covering perception, spatial and quantitative reasoning, compliance assessment, evidence synthesis, causal analysis, and mitigation-oriented decision making. Evaluation of representative proprietary and open-source vision--language models reveals substantial performance differences and persistent weaknesses in comparative, technical, and multi-evidence reasoning, demonstrating that strong general visual understanding does not yet guarantee reliable industrial-safety reasoning.
Problem

Research questions and friction points this paper is trying to address.

industrial safety
multimodal reasoning
accident knowledge
evidence-grounded reasoning
safety compliance
Innovation

Methods, ideas, or system contributions that make the work stand out.

multimodal reasoning
industrial safety
scene graph
evidence-based QA
accident knowledge integration
🔎 Similar Papers
2024-10-08arXiv.orgCitations: 3
2024-07-082024 IEEE International Automated Vehicle Validation Conference (IAVVC)Citations: 1
Y
Yuanchi Zhu
ShanghaiTech University
K
Kang An
Shanghai Jiao Tong University
T
Tengyue Wang
South China University of Technology
Z
Zhongyu Yang
ModelBest
C
Chenxu Du
Southwest Jiaotong University
X
Xinqi Yang
East China Normal University
H
Hebao Zhu
Chongqing University
B
Bokai Zhao
University of the Chinese Academy of Sciences
T
Tianyu Liang
Southeast University
Z
Ziliang Wang
SenseTime
F
Faqiang Qian
SenseTime
Y
Yunli Yang
Institute for Advanced Algorithms Research, Shanghai
W
Weiyang Shi
Institute of Automation, Chinese Academy of Sciences
Qibing Ren
Qibing Ren
Shanghai Jiao Tong University
machine learningcomputer visiontrustworthy AI