LongDocBench: Benchmarking TOC Hierarchy and Contextual Relationship Recovery in Long Documents

📅 2026-08-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of evaluation for table-of-contents hierarchies and contextual relationships in existing benchmarks by constructing a novel dataset comprising 85 real-world documents alongside a dedicated evaluation protocol. By providing human-verified structural annotations and designing structured parsing experiments, we systematically assess model capabilities in structure recovery. Our findings demonstrate that accurate structural restoration significantly enhances reasoning performance in long-document question answering. Furthermore, this work reveals critical limitations of current parsers in document-level tasks, thereby establishing essential infrastructure and offering new insights for advancing long-document understanding research.
📝 Abstract
Parsing visual documents into machine-readable representations is fundamental to document intelligence. Existing benchmarks focus on page-level element recognition, reading order, formula recognition, and table structure. Long documents, however, also require document-level structure recovery. This includes reconstructing cross-page table-of-contents (TOC) hierarchies and identifying typed links from tables and figures to their captions, notes, and sources, often in one-to-many form. Because these structures are covered only partially or subsumed within broader parsing protocols, existing benchmarks cannot directly evaluate two key document-level tasks: \emph{Table-of-Contents Hierarchy Recovery} and \emph{Contextual Relationship Recovery}. To benchmark these two tasks, we introduce \textsc{LongDocBench}, comprising 85 real-world financial reports, textbooks, and academic papers spanning 2,582 pages, with up to 105 pages per document. It provides human-verified annotations for 3,937 heading nodes (mean node depth 3.55; maximum depth 9) and 3,258 contextual relationships annotated across 2,680 table and figure objects. We further evaluate both the downstream utility and recoverability of these structures. Long-document question-answering experiments show that human-verified TOC hierarchies and contextual relationships improve reasoning, with their combination providing complementary benefits. Meanwhile, representative document parsers remain limited on both recovery tasks despite strong page-level performance. To support further progress, we publicly release \textsc{LongDocBench} and its evaluation protocol and reproducible testbed for advancing document-level structure recovery in long documents.
Problem

Research questions and friction points this paper is trying to address.

Long Document Parsing
TOC Hierarchy Recovery
Contextual Relationship Recovery
Document Intelligence Benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

LongDocBench
TOC Hierarchy Recovery
Contextual Relationship Recovery
Document-level Structure
Long Document Parsing
💼 Related Jobs
No related jobs found.
Y
Yuefeng Zou
Ant Group
Y
Yichen Lu
Ant Group
J
Jingxiao Yang
Ant Group, Zhejiang University
Bingtao Fu
Bingtao Fu
Ant Group
Computer Vision
G
Gaoyang Zhang
Ant Group
X
Xiongfei Bai
Ant Group
T
Tian Chen
Ant Group
X
Xiang Qi
Ant Group