Beyond Blind Compliance: Benchmarking Task Verification in OCR Reasoning

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究针对OCR场景中任务验证问题,通过构建包含1800个样本的VeriOCRBench基准,评估多模态大语言模型在任务验证上的表现,揭示了现有系统的可靠性差距。
📝 Abstract
Multimodal Large Language Models (MLLMs) have achieved strong performance on OCR-centric document understanding and text-rich visual reasoning benchmarks. Yet existing evaluations largely assume that every task is valid and answerable. In real-world OCR scenarios, this assumption often fails: questions may rely on illegible text, occluded evidence, nonexistent visual targets, contradictory premises, or missing variables. We study this reliability gap as OCR-grounded Task Verification: before answering, a model should determine whether the Image Premise (IP), Textual Premise (TP), and Question (Q) jointly define an executable task. We introduce VeriOCRBench, a 1,800-sample human-verified benchmark built from source images drawn from 8 OCR-related datasets and spanning 8 real-world image domains, with controlled, image-grounded diagnostic tasks. It contains 1,600 trap-injected invalid tasks across 8 trap types and four verification dimensions---Visual, Contextual, Factual, and Logical---plus 200 trap-free controls for measuring over-refusal. Built with a Visual Atomic Fact (VAF)-anchored pipeline and full human auditing, VeriOCRBench enables decoupled evaluation of task verification, root-cause diagnosis, and over-refusal. Evaluating 15 leading MLLMs reveals persistent blind compliance, diagnosis failures, and prompt-induced over-refusal, exposing a critical reliability gap in current OCR reasoning systems. The code is available at: https://github.com/zy001122/Beyond-Blind-Compliance.
Problem

Research questions and friction points this paper is trying to address.

OCR
Task Verification
Multimodal Large Language Models
Document Understanding
Visual Reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Task Verification
VeriOCRBench
Multimodal Large Language Models
OCR Reasoning
🔎 Similar Papers
No similar papers found.
Y
Yue Zhou
School of Artificial Intelligence, Jilin University
Y
Yuan Wu
School of Artificial Intelligence, Jilin University
Yi Chang
Yi Chang
Jilin University
Information RetrievalData MiningNatural Language ProcessingMachine LearningArtificial Intelligence