Evaluating Structured Information Extraction with Open Models in a High Risk Public Sector Application

📅 2026-08-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究针对高风险公共部门应用中的信息提取问题,通过评估开源OCR引擎、大语言模型和视觉-语言模型在处理复杂文档任务上的表现,揭示了现有模型的局限性和影响因素。
📝 Abstract
The extraction of structured information from unstructured documents represents a critical component of digital transformations in all sectors. While proprietary solutions dominate commercial applications, a rapidly growing ecosystem of open-source Optical Character Recognition (OCR) engines, Large Language Models (LLMs), and Vision-Language Models (VLMs) offers accessible alternatives. However, systematic evaluations on realistic, multi-step extraction pipelines remain scarce. Responsible usage of such extraction tools require comprehensive evaluations on realistic tasks, especially as these solutions will be key components of applications in the public sector that the EU AI act categorizes as high risk. To address this gap we present a comprehensive benchmark assessing the end-to-end performance of open-source systems on a complex real-world document processing task classified as high risk: Student applications for an international study program. We conduct a comprehensive empirical evaluation with state-of-the-art OCR engines, LLMs and VLMs. Our results reveal that while VLMs generally outperform OCR+LLM pipelines, even state-of-the-art open-source models struggle to handle such tasks reliably in zero-shot settings. Only 4 of 35 configurations achieved F1 scores above 0.5, with the best OCR+LLM pipeline matching top VLM performance, though most OCR+LLM combinations performed substantially worse. Roughly 75\% of all configurations scored below 0.25. Model scale influences performance, yet the relationship is non-linear: substantially larger models do not guarantee proportionally better results. Input quality, particularly the structural preservation of OCR output, emerges as a critical factor independent of downstream model capability.
Problem

Research questions and friction points this paper is trying to address.

Structured Information Extraction
Open-source Models
High Risk Public Sector
End-to-End Performance
Realistic Document Processing
Innovation

Methods, ideas, or system contributions that make the work stand out.

Benchmarking
Open-source Systems
High-risk Applications
VLMs vs. OCR+LLM
Input Quality
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
E
Elias Schubert
Berlin University of Applied Sciences, Luxemburger Str. 10, 13353 Berlin, Germany
F
Felix Bießmann
Berlin University of Applied Sciences, Luxemburger Str. 10, 13353 Berlin, Germany; Einstein Center Digital Future, Wilhelmstraße 67, 10117 Berlin, Germany