Robust Diagram Reasoning: A Framework for Enhancing LVLM Performance on Visually Perturbed Scientific Diagrams

📅 2025-08-23
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Current large vision-language models (LVLMs) exhibit severe robustness deficiencies in understanding scientific charts under realistic visual perturbations—such as noise, blur, and occlusion—while existing evaluation benchmarks lack systematic assessment of such perturbations. Method: We propose the first programmatic perturbation-robustness evaluation framework for scientific chart understanding: (1) constructing a large-scale scientific chart QA dataset with diverse, controllable perturbations; (2) designing an adaptive multi-view reasoning mechanism with consistency-based self-correction; and (3) introducing two novel quantitative metrics—Perturbation Robustness Consistency (PRC) and Visual Degradation Sensitivity (VDC). Results: Experiments demonstrate substantial improvements in LVLM reasoning performance across perturbation types. Notably, mainstream models—including GPT-4V—suffer over 13 percentage-point accuracy drops under perturbation, revealing critical robustness gaps. Our framework establishes a new paradigm for evaluating and advancing robustness in scientific multimodal models.

Technology Category

Application Category

📝 Abstract
Large Language Models (LLMs) and their multimodal variants (LVLMs) hold immense promise for scientific and engineering applications, particularly in processing visual information like scientific diagrams. However, their practical deployment is hindered by a critical lack of robustness to common visual perturbations such as noise, blur, and occlusions, which are prevalent in real-world scientific documents. Existing evaluation benchmarks largely overlook this challenge, leaving the robust reasoning capabilities of LVLMs on visually degraded scientific diagrams underexplored. To address this, we introduce the Robust Diagram Reasoning (RDR) framework, a novel approach designed to enhance and rigorously evaluate LVLMs' performance under such conditions. At its core, RDR employs an Adaptive Multi-View & Consistency Verification (AMCV) mechanism, which involves generating multiple perturbed versions of a diagram, performing parallel inference, and then applying a consistency-based self-correction loop. We also propose two new metrics, Perturbation Robustness Score (PRS) and Visual Degradation Consistency (VDC), to quantify robustness. Furthermore, we construct SciDiagram-Robust, the first large-scale scientific diagram question-answering dataset specifically augmented with diverse, programmatically generated visual perturbations. Our extensive experiments demonstrate that even state-of-the-art closed-source LVLMs like GPT-4V exhibit significant performance degradation when faced with perturbed inputs (Clean Accuracy 85.2% vs. PRS 72.1%).
Problem

Research questions and friction points this paper is trying to address.

Addressing LVLMs' lack of robustness to visual perturbations in diagrams
Evaluating model performance on degraded scientific diagrams with noise and blur
Proposing a framework to enhance reasoning consistency under visual corruption
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adaptive Multi-View & Consistency Verification mechanism
Generating multiple perturbed diagram versions
Consistency-based self-correction loop
🔎 Similar Papers
M
Minghao Zhou
Taiyuan University of Science and Technology
R
Rafael Souza
University of Brasilia
Y
Yaqian Hu
Taiyuan University of Science and Technology
L
Luming Che
Taiyuan University of Science and Technology