Rethinking the Evaluation of Pre-trained Text-and-Layout Models from an Entity-Centric Perspective

πŸ“… 2024-02-04
πŸ›οΈ arXiv.org
πŸ“ˆ Citations: 1
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Existing evaluation benchmarks for vision-rich document (VrD) information extraction suffer from severe biases: coarse-grained annotations introduce spurious correlations between inputs and labels, inflating model performance estimates; moreover, holistic F1-based evaluation fails to assess robustness in realistic scenarios. Method: We propose an entity-centric evaluation paradigm and introduce EC-FUNSDβ€”the first semantic-driven benchmark for entity recognition and linking in VrDs. Contribution/Results: EC-FUNSD innovatively decouples paragraph-block layout annotations from semantic entity definitions, incorporating cross-block entity linking, fine-grained relational understanding, and diverse document layouts. Experiments show that state-of-the-art vision-language pre-trained models exhibit substantial performance degradation on EC-FUNSD, validating the overfitting issue inherent in prior benchmarks. EC-FUNSD thus establishes a more rigorous, semantically grounded, and robust evaluation standard for VrD understanding.

Technology Category

Application Category

πŸ“ Abstract
Recently developed pre-trained text-and-layout models (PTLMs) have shown remarkable success in multiple information extraction tasks on visually-rich documents. However, the prevailing evaluation pipeline may not be sufficiently robust for assessing the information extraction ability of PTLMs, due to inadequate annotations within the benchmarks. Therefore, we claim the necessary standards for an ideal benchmark to evaluate the information extraction ability of PTLMs. We then introduce EC-FUNSD, an entity-centric benckmark designed for the evaluation of semantic entity recognition and entity linking on visually-rich documents. This dataset contains diverse formats of document layouts and annotations of semantic-driven entities and their relations. Moreover, this dataset disentangles the falsely coupled annotation of segment and entity that arises from the block-level annotation of FUNSD. Experiment results demonstrate that state-of-the-art PTLMs exhibit overfitting tendencies on the prevailing benchmarks, as their performance sharply decrease when the dataset bias is removed.
Problem

Research questions and friction points this paper is trying to address.

PTLMs' real-world performance lags behind benchmark results
Benchmark datasets contain inadequate annotations causing biased evaluations
Current evaluation methods fail to assess real-world capabilities comprehensively
Innovation

Methods, ideas, or system contributions that make the work stand out.

Introduces EC-FUNSD dataset for VrD information extraction
Disentangles falsely-coupled segment and entity annotations
Evaluates PTLMs on generalization, robustness, and fairness
πŸ”Ž Similar Papers
No similar papers found.
Fudan University | Nankai University | Ant Group Inc.
C
Chong Zhang
School of Computer Science, Fudan University, Shanghai, China
Y
Yixi Zhao
School of Computer Science, Fudan University, Shanghai, China
C
Chenshu Yuan
School of Statistics and Data Science, Nankai University, Tianjin, China
Yi Tu
Yi Tu
Ant Group
Computer VisionDocument UnderstandingVision Language Model
Y
Ya Guo
Ant Info Security Lab, Ant Group Inc., Hangzhou, China
Q
Qi Zhang
School of Computer Science, Fudan University, Shanghai, China