GDP.pdf: Benchmarking Grounded Multimodal Reasoning over Professional PDF Documents

📅 2026-07-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Current document-centric AI benchmarks often assess individual capabilities in isolation, failing to capture models’ holistic reasoning performance on PDFs in authentic professional contexts. To address this gap, this work introduces a high-quality benchmark comprising question–document pairs authored by experts across ten professional domains, retaining only those samples on which at least two state-of-the-art multimodal models commit substantive errors. The study uniquely focuses on real-world querying scenarios involving professional PDFs and proposes an atomic-criteria–based scoring mechanism alongside a three-tier, eleven-category capability taxonomy. It also emphasizes the model’s ability to abstain when queries lack sufficient support in the document. Evaluation reveals that even the best-performing model passes only 15% of the 100 test cases, with primary failure modes including table misalignment, chart misinterpretation, footnote omission, symbol-counting errors, and mishandling of revised or overlaid text.
📝 Abstract
A large share of day-to-day work in professional domains happens inside PDF files: benefits packets, leases, datasheets, clinical guidelines, construction plans. Benchmarks for document AI have generally measured the required capabilities in isolation: OCR, layout analysis, chart reasoning, table QA, document VQA. A high score on any one of them does not necessarily reveal whether a model can answer a realistic question that someone in the field would actually ask about a specific PDF. GDP.pdf is a benchmark built to measure this directly. It consists of question-document pairs authored by working professionals in ten fields, and a candidate question was kept only when at least two frontier multimodal models failed it in a way that mattered: a wrong answer, missed decisive evidence, or a fabricated claim, rather than a superficial difference such as style. Each item comes with a rubric of atomic criteria, so we can report a graded rubric score as well as a strict task-level pass rate, and each item is tagged against a taxonomy of eleven capabilities in three tiers, spanning text extraction and grounding, table and chart comprehension, cross-referencing, spatial reasoning, and abstention on unsupported queries. We evaluated seven frontier models on the 100-item benchmark. The best model passed only 15% of the items and the worst passed 1%. Most errors trace back to a small set of recurring loss patterns: misaligned tables, misread charts, skipped footnotes and exclusions, miscounted floor-plan symbols, scan noise, and amendments that supersede earlier text. The full 100-item benchmark is publicly available at https://huggingface.co/datasets/surgeai/GDP.pdf
Problem

Research questions and friction points this paper is trying to address.

document AI
multimodal reasoning
professional PDF
benchmarking
grounded reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

grounded multimodal reasoning
professional PDF benchmark
rubric-based evaluation
capability taxonomy
document AI failure analysis
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
S
Suhaas Garre
Surge AI
E
Emily Ritchie
Surge AI
S
Sushant Mehta
Surge AI
E
Edwin Chen
Surge AI