Institution profile

China Mobile Information Technology Co. Ltd.

Industry researchasia · cn
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Visual Information Extraction from Documents via Classification-Guided Large Vision-Language Models

Jul 21, 2026

This work addresses the challenge of information extraction from visually rich documents, where diverse layouts and real-world interferences—such as stamps and low contrast—hinder performance. Existing approaches typically rely on extensive annotated data and layout-specific training. To overcome these limitations, the authors propose a classification-guided large vision-language model framework that decouples document-type classification from content extraction and leverages in-context learning to dynamically construct prompts, enabling zero-shot information extraction across multiple document types. Built upon Qwen2.5-VL-7B, the method employs conditional computation to reduce task uncertainty. Evaluated on a real-world tendering dataset comprising 16 certificate categories, it achieves a zero-shot F1 score of 86.43%, outperforming a strong supervised baseline by 18.35 percentage points; with fine-tuning, F1 further improves to 93.65%, attaining a normalized edit distance of 0.93.

0 citationsRead paper
Recent publications

Latest Papers

Visual Information Extraction from Documents via Classification-Guided Large Vision-Language Models

Jul 21, 2026

This work addresses the challenge of information extraction from visually rich documents, where diverse layouts and real-world interferences—such as stamps and low contrast—hinder performance. Existing approaches typically rely on extensive annotated data and layout-specific training. To overcome these limitations, the authors propose a classification-guided large vision-language model framework that decouples document-type classification from content extraction and leverages in-context learning to dynamically construct prompts, enabling zero-shot information extraction across multiple document types. Built upon Qwen2.5-VL-7B, the method employs conditional computation to reduce task uncertainty. Evaluated on a real-world tendering dataset comprising 16 certificate categories, it achieves a zero-shot F1 score of 86.43%, outperforming a strong supervised baseline by 18.35 percentage points; with fine-tuning, F1 further improves to 93.65%, attaining a normalized edit distance of 0.93.

0 citationsRead paper