WildHandBench: A Benchmark for Handwritten Text Understanding that Challenges MLLMs and Humans

📅 2026-08-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过创建WildHandBench基准,解决了现有模型在处理手写文档时能力不足的问题,并引入了Prior-Driven Error指标来分析错误来源。
📝 Abstract
While the top model on OmniDocBench now reaches 96.34% overall on printed-document parsing, the ability of current models to handle challenging handwritten documents remains largely uncharacterized. Existing benchmarks focus on isolated text or formulas, overlook handwritten tables and real-world degradation, and report aggregate accuracy without explaining why models fail. We present WildHandBench, a benchmark containing 500 handwritten documents across three structures (free text, tables, formulas), four languages, and nine real-world scenarios. We introduce a Prior-Driven Error (PDE) metric that quantifies whether errors originate from language priors rather than visual evidence. Evaluating 18 state-of-the-art models together with calibrated human baselines, we find: (1) the best model achieves only 71.85% overall; (2) humans outperform all models yet the gap is narrow (77.09% vs. 71.85%); and (3) model errors are qualitatively different from human errors -- 63-91% of model errors are prior-driven versus only 49% for humans, exposing systematic reliance on language priors that conventional accuracy metrics cannot capture.
Problem

Research questions and friction points this paper is trying to address.

handwritten documents
benchmark
model errors
language priors
visual evidence
Innovation

Methods, ideas, or system contributions that make the work stand out.

WildHandBench
Prior-Driven Error (PDE)
handwritten document understanding
💼 Related Jobs
No related jobs found.
J
Jun Zhang
Baidu Inc.
Q
Qiao Zhao
Baidu Inc.
Cheng Cui
Cheng Cui
BUAA
deep learningnetwork designOCRmllm
J
Jianying Qu
Baidu Inc.
Zhongkai Sun
Zhongkai Sun
Amazon Alexa AI
J
Jianwen Yang
Baidu Inc.
C
Changda Zhou
Baidu Inc.
Z
ZhuoXin Liu
Baidu Inc.
S
Shubin Han
Baidu Inc.