DocIntent: Answerability-Guided Agentic Restoration for Real-World Document Visual Question Answering

📅 2026-08-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
针对文档视觉问答中现实世界退化问题,提出DocIntent框架,通过评估问题可回答性指导自适应恢复过程,无需额外预训练模型,提升多模态大语言模型性能。
📝 Abstract
Real-world degradations such as blur, shadow, distortion, and moire patterns severely impair the document question-answering capabilities of Multimodal Large Language Models (MLLMs). Applying restoration tools before Visual Question Answering (VQA) is an intuitive solution. However, existing restoration approaches remain limited, as manually designing and executing restoration strategies is labor-intensive and requires domain expertise. Agentic restoration offers new possibilities for automation, yet existing frameworks primarily target natural images and pursue perceptual quality, overlooking that restoration should serve downstream tasks rather than optimize generic image quality metrics. To this end, we explore the value of agentic restoration for real-world degraded document VQA and propose DocIntent, a training-free Answerability-Guided Agentic Restoration framework. DocIntent first assesses question answerability, then identifies task-relevant degradations and selectively invokes restoration tools. A Comparison-Based Rollback mechanism validates each restoration step and reverts it when question-relevant evidence becomes less decipherable. The entire process requires no additional pretrained degradation classifier or image quality assessment model. Extensive experiments on the WildDoc benchmark show that DocIntent consistently improves the average score and consistency of different open- and closed-source MLLMs. The code and experimental data will be publicly available.
Problem

Research questions and friction points this paper is trying to address.

Document Visual Question Answering
Real-World Degradations
Agentic Restoration
Answerability-Guided
Innovation

Methods, ideas, or system contributions that make the work stand out.

Answerability-Guided Agentic Restoration
Comparison-Based Rollback
Real-World Document VQA
🔎 Similar Papers
2024-07-17European Conference on Computer VisionCitations: 2
💼 Related Jobs
No related jobs found.
Z
Zihan Huang
South China University of Technology
S
Shihang Wu
South China University of Technology
Junle Liu
Junle Liu
South China University of Technology
AIGC
P
Peirong Zhang
South China University of Technology
Yongxin Shi
Yongxin Shi
South China University of Technology
Computer VisionOCRMultimodal LLMs
X
Xuhan Zheng
South China University of Technology
Lianwen Jin
Lianwen Jin
Professor of Electronic and Information Engineering, South China University of Technology
Optical Character Recognition (OCR)Computer VisionDocument AIMultimodal LLMs