ReViCo: Unveiling the Limitations of VLMs in Visual Text Understanding via Error Correction

📅 2026-08-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过引入ReViCo基准,旨在评估视觉语言模型在图像中文本理解上的局限性,并通过错误纠正任务挑战这些模型,揭示了当前模型与人类表现之间的显著差距。
📝 Abstract
Vision Language Models (VLMs) have shown great success in general visual tasks, yet they still struggle to deeply understand text within images. In this paper, we introduce ReViCo (Real Visual Correction), a benchmark designed to evaluate VLM text understanding through a novel task of visual text error correction. ReViCo challenges models to identify and fix text errors in real-world images, which requires a profound understanding of the interplay between visual text and its surrounding visual context. We benchmark various VLMs using two distinct paradigms: prompt-based strategy and targeted model training, both aimed at pushing the limits of current models. Our experiments reveal a striking performance gap between even the best VLMs and human, and further analysis also shows that most models struggle to accurately perceive the visual text, resulting in frequent correction errors. By highlighting these gaps, ReViCo provides a new benchmark foundation for developing more robust and text-aware VLMs.
Problem

Research questions and friction points this paper is trying to address.

Vision Language Models
text understanding
visual text error correction
visual context
Innovation

Methods, ideas, or system contributions that make the work stand out.

Visual Text Understanding
Error Correction
Benchmark
Vision Language Models
Contextual Awareness
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
B
Bojun Zhang
State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, CAS, Beijing, China; School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China
J
Junhong Liang
Mohamed bin Zayed University of Artificial Intelligence
Feifei Zhai
Feifei Zhai
Institute of Automation, Chinese Academy of Sciences
Machine TranslationNatural Language ProcessingMachine Learning
Fengxian Ji
Fengxian Ji
Northeast University
agent、Machine learnin、CV
Y
Yu Zhou
State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, CAS, Beijing, China; Fanyu AI Laboratory, Zhongke Fanyu Technology Co., Ltd, Beijing, China