BRUCE: Benchmarking Robustness Under Corruption Escalation for Scientific Vision-Language Reasoning

📅 2026-08-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the significant degradation in scientific reasoning capabilities of current vision-language models (VLMs) under visual corruptions such as blurriness and low contrast. To this end, we propose BRUCE, a framework that systematically evaluates multimodal model robustness against visual degradations for the first time. By applying diverse image perturbations and assessing performance across four reasoning dimensions—OCR dependency, spatial, symbolic, and semantic—we conduct fine-grained evaluations on chemistry and mathematics tasks. We further introduce two novel metrics, RCI and T-RCI, to quantify the rate at which reasoning performance declines with increasing visual corruption. Our experiments reveal the high sensitivity of state-of-the-art VLMs to visual corruptions and establish an interpretable taxonomy of failure modes.
📝 Abstract
Visual-language models (VLMs) frequently struggle with robustness issues in real-world situations due to low- or varying-quality input images. In this paper, we aim at analyzing VLMs' robustness by applying perturbations and distortions to the input images, such as blur or low contrast. Toward this goal, we propose BRUCE (Benchmarking Robustness Under Corruption Escalation, a multimodal reasoning fragility framework for scientific vision-language reasoning. State-of-the-art evaluation frameworks/studies primarily focus on clean-task accuracy and rarely analyze how reasoning stability degrades across robustness dimensions. Besides varying over a wide-range of input perturbations, BRUCE employs two novel metrics -- Robustness Corruption Index (RCI) and Traversal-RCI (T-RCI) -- to quantify how rapidly multimodal reasoning performance deteriorates in VLMs as visual corruption severity increases under progressive perturbation scaling. We evaluate BRUCE across chemistry and mathematical reasoning tasks for multiple datasets, while analyzing corruption-induced prediction failures in terms of four high-level reasoning domains: OCR-dependent reasoning, spatial reasoning, symbolic reasoning, and semantic failures, with each containing fine-grained corruption specific failure subtypes, thereby enabling an interpretable failure analysis.
Problem

Research questions and friction points this paper is trying to address.

robustness
vision-language models
input corruption
reasoning stability
multimodal reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Robustness Benchmarking
Vision-Language Models
Corruption Escalation
RCI Metric
Multimodal Reasoning