DocIQ: A Benchmark Dataset and Feature Fusion Network for Document Image Quality Assessment

πŸ“… 2025-09-21
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Document image quality assessment (DIQA) suffers from a lack of large-scale, high-quality subjective datasets and dedicated no-reference models. Method: We introduce DIQA-5000β€”the first large-scale subjective DIQA dataset comprising 5,000 document imagesβ€”and propose a lightweight no-reference model that jointly leverages multi-level visual features and document layout priors. Our approach innovates with a layout-aware downsampling mechanism to preserve structural sensitivity at low resolutions, a multi-quality-head architecture to separately model fine-grained distributions of overall quality, sharpness, and color fidelity, and a feature fusion module to synergistically optimize low-level texture and high-level semantic representations. Contribution/Results: Extensive experiments demonstrate that our method significantly outperforms state-of-the-art general-purpose image quality assessment (IQA) models on both DIQA-5000 and OCR-oriented benchmarks, validating its effectiveness, generalizability, and practical utility for real-world document processing tasks.

Technology Category

Application Category

πŸ“ Abstract
Document image quality assessment (DIQA) is an important component for various applications, including optical character recognition (OCR), document restoration, and the evaluation of document image processing systems. In this paper, we introduce a subjective DIQA dataset DIQA-5000. The DIQA-5000 dataset comprises 5,000 document images, generated by applying multiple document enhancement techniques to 500 real-world images with diverse distortions. Each enhanced image was rated by 15 subjects across three rating dimensions: overall quality, sharpness, and color fidelity. Furthermore, we propose a specialized no-reference DIQA model that exploits document layout features to maintain quality perception at reduced resolutions to lower computational cost. Recognizing that image quality is influenced by both low-level and high-level visual features, we designed a feature fusion module to extract and integrate multi-level features from document images. To generate multi-dimensional scores, our model employs independent quality heads for each dimension to predict score distributions, allowing it to learn distinct aspects of document image quality. Experimental results demonstrate that our method outperforms current state-of-the-art general-purpose IQA models on both DIQA-5000 and an additional document image dataset focused on OCR accuracy.
Problem

Research questions and friction points this paper is trying to address.

Assessing document image quality for OCR and restoration applications
Creating a benchmark dataset with subjective quality ratings
Developing feature fusion model for multi-dimensional quality assessment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Feature fusion module integrates multi-level visual features
Independent quality heads predict multi-dimensional score distributions
Document layout features maintain quality at reduced resolutions
πŸ’Ό Related Jobs
No related jobs found.
Z
Zhichao Ma
Shanghai Jiao Tong University, Shanghai, China
F
Fan Huang
Shanghai Jiao Tong University, Shanghai, China
L
Lu Zhao
INTSIG Information Co. Ltd, Shanghai, China
F
Fengjun Guo
INTSIG Information Co. Ltd, Shanghai, China
Guangtao Zhai
Guangtao Zhai
Professor, IEEE Fellow, Shanghai Jiao Tong University
Multimedia Signal ProcessingVisual Quality AssessmentQoEAI EvaluationDisplays
X
Xiongkuo Min
Shanghai Jiao Tong University, Shanghai, China