Identify, Locate, Link: End-to-End Key-Value Extraction from Document Images

📅 2026-08-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过微调SmolDocling模型直接从文档图像中端到端提取键值对,解决了传统流程中的多阶段错误传播问题。
📝 Abstract
Document processing pipelines traditionally cascade optical character recognition (OCR) engines with downstream models for structured information extraction, leading to multi-stage error propagation. We fine-tune SmolDocling, a compact 256M-parameter vision-language model (VLM), to perform end-to-end key-value extraction directly from document images, jointly solving identification, localization, and association in a single pass without OCR preprocessing. We extend DocTags with specialized key, value, region, and link tags, enabling many-to-many relationships in a unified output sequence. To address data limitations, we design an augmentation pipeline combining synthetic form filling and graph-based crops that preserve complete key-value subgraphs. We further introduce a layout-aware evaluation framework extending text matching with spatial bounding box verification. On FUNSD, XFUND, and a large-scale private dataset, our model outperforms larger zero-shot VLM baselines under layout-aware evaluation, while being 27 times smaller than Qwen2.5-VL (7B) and over 5 times faster at inference. The model weights will be released publicly after publication.
Problem

Research questions and friction points this paper is trying to address.

end-to-end key-value extraction
document images
multi-stage error propagation
identification and localization
Innovation

Methods, ideas, or system contributions that make the work stand out.

end-to-end key-value extraction
vision-language model
layout-aware evaluation
data augmentation
💼 Related Jobs
No related jobs found.
A
A. Said Gurbuz
IBM Research Zurich, Rüschlikon, Switzerland; ETH Zurich, Zurich, Switzerland
A
Ahmed Nassar
IBM Research Zurich, Rüschlikon, Switzerland
Christoph Auer
Christoph Auer
IBM Research
Maksym Lysak
Maksym Lysak
IBM
artificial intelligencecomputer vision3d graphicsartshistory
L
Lucas Morin
IBM Research Zurich, Rüschlikon, Switzerland
M
Matteo Omenetti
IBM Research Zurich, Rüschlikon, Switzerland
T
Tim Strohmeyer
IBM Research Zurich, Rüschlikon, Switzerland
P
Panagiotis Vagenas
IBM Research Zurich, Rüschlikon, Switzerland
Nikolaos Livathinos
Nikolaos Livathinos
IBM Research
Computer VisionAISoftware Architecture
Michele Dolfi
Michele Dolfi
IBM Research
Knowledge ingestionCloud computingComputational physicsTensor networksHigh performance computing
P
Peter Staar
IBM Research Zurich, Rüschlikon, Switzerland