Towards reliable multimodal disaster severity assessment through preference optimization and explainable vision-language reasoning

📅 2026-09-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对灾害评估中数据稀缺和推理质量不足的问题,提出结合监督微调和直接偏好优化的两阶段训练框架,提升预测准确性和解释透明度。
📝 Abstract
Reliable disaster damage assessment requires models that provide both accurate predictions and transparent explanations. However, existing multimodal approaches are limited by scarce annotated data and insufficient evaluation of reasoning quality. This study proposes a two-stage training framework that integrates Supervised Fine-Tuning (SFT) and Direct Preference Optimization (DPO) within a unified data construction pipeline. From a single Human-in-the-Loop (HITL) annotation workflow, two complementary datasets are derived, namely ReasoningSet, which contains validated rationales for SFT, and PreferenceSet, which comprises paired rationales for DPO-based alignment. The framework evaluates both classification performance and explanation quality using automatic metrics, model-based scoring, and human ranking. Experimental results show that SFT improves accuracy from 73.64% to 78.29% and increases Macro-F1 by 29% compared to the baseline, while explanation quality improves by approximately 25%. Subsequent DPO alignment further enhances interpretability on the PreferenceSet. Cross-model validation on InternVL-3-8B and LLaVA-1.5-7B demonstrates the robustness and generalizability of the approach. The proposed framework improves detection of underrepresented mild damage cases, reduces high-risk misclassifications, and strengthens alignment between model reasoning and human judgment. Overall, it provides a reproducible pathway to develop reliable multimodal systems that deliver auditable, actionable disaster insights for emergency management.
Problem

Research questions and friction points this paper is trying to address.

disaster damage assessment
multimodal approaches
annotated data
reasoning quality
Innovation

Methods, ideas, or system contributions that make the work stand out.

Supervised Fine-Tuning (SFT)
Direct Preference Optimization (DPO)
ReasoningSet
PreferenceSet
Human-in-the-Loop (HITL)
💼 Related Jobs
No related jobs found.
Y
Yuanjun Zhang
Centre for Machine Vision & Signal Processing, University of Oulu, Pentti Kaiteran katu 1, PO Box 8000, Oulu, 90570, Finland
Fuzel Ahamed Shaik
Fuzel Ahamed Shaik
University of Oulu
MLDLNLPData Science
S
Suvojit Acharjee
Department of CSE and Institute of Engineering and Management, 120 SDF Building, Saltlake Electronics Complex, Kolkata, 700091, India
F
Fahad Khalid
Centre for Machine Vision & Signal Processing, University of Oulu, Pentti Kaiteran katu 1, PO Box 8000, Oulu, 90570, Finland
Mourad Oussalah
Mourad Oussalah
University of Oulu
Social mediadata miningroboticsdata fusioncomputer vision