Towards Efficient Multimodal and Multilingual Opinion Extraction for STI: A QLoRA-Based Fine-Tuning Approach

📅 2026-08-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses noise interference and unreliable zero-shot structured output in multimodal opinion extraction for scientific and technological intelligence. We propose a vision-anchored core opinion extraction framework that fine-tunes VideoLLaMA2 via QLoRA and incorporates a visual evidence anchoring mechanism. Furthermore, fuzzy cumulative prospect theory is integrated for value assessment and post-processing to enable efficient structured filtering. Experimental results demonstrate that the model achieves an F1 score of 51.14% and an accuracy of 74%. Notably, F1 scores for Spanish and Russian increase dramatically from 4.83% and 0.45% to 46.05% and 51.93%, respectively, significantly enhancing structured extraction capabilities for low-resource languages in intelligence analysis.
📝 Abstract
Recent advances in large language models (LLMs) have reshaped semantic analysis. Opinion Extraction (OE) for Science and Technology Intelligence (STI) requires concise core opinions from large information streams. Off-the-shelf models struggle to filter noise from these streams and show limited structured-output reliability in zero-shot multilingual and multi-modal settings. To address information overload and extraction defocus, this study proposes a multimodal core-opinion extraction framework in which visual evidence serves as a contextual anchor for textual judgment. Using VideoLLaMA2 (VL2) and VideoLLaMA2.1 (VL2.1) as the base models, we apply Quantized Low-Rank Adaptation (QLoRA) fine-tuning on a curated dataset of 2,194 multilingual and multimodal samples. Under the selected Image-Augmented setting, fine-tuned VL2.1 generates structured JSON core-opinion outputs, achieving 64.98% Precision, 42.15% Recall, 51.14% F1-score, and 74.00% sample-level accuracy. Relative to the zero-shot VL2.1 setting, it raises the F1-scores of Spanish and Russian from 4.83% and 0.45% to 46.05% and 51.93%, respectively. The framework further incorporates a Fuzzy Cumulative Prospect Theory-based post-extraction triage module for case-level value assessment, providing a case-level value signal for downstream STI screening.
Problem

Research questions and friction points this paper is trying to address.

Opinion Extraction
Science and Technology Intelligence
Multimodal
Multilingual
Large Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

QLoRA Fine-Tuning
Multimodal Opinion Extraction
Visual Contextual Anchor
Fuzzy Cumulative Prospect Theory
Structured Output Generation
🔎 Similar Papers
No similar papers found.
S
Sheng Hong
School of Cyber Science and Technology, Beihang University, Beijing, 100191, China
X
Xuanqi Wang
School of Information and Engineering, Nanchang University, Nanchang, 330031, China
Jiacheng Wang
Jiacheng Wang
Nanyang Technological University
ISACGenAILow-altitude wireless networkSemantic Communications
Y
Yuwei Wang
Institute of Computing Technology, Chinese Academy of Sciences, Beijing, 100190, China