Institution profile

NingboTech University

Academic institutionasia · cn
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

Visual Information Extraction from Documents via Classification-Guided Large Vision-Language Models

Jul 21, 2026

This work addresses the challenge of information extraction from visually rich documents, where diverse layouts and real-world interferences—such as stamps and low contrast—hinder performance. Existing approaches typically rely on extensive annotated data and layout-specific training. To overcome these limitations, the authors propose a classification-guided large vision-language model framework that decouples document-type classification from content extraction and leverages in-context learning to dynamically construct prompts, enabling zero-shot information extraction across multiple document types. Built upon Qwen2.5-VL-7B, the method employs conditional computation to reduce task uncertainty. Evaluated on a real-world tendering dataset comprising 16 certificate categories, it achieves a zero-shot F1 score of 86.43%, outperforming a strong supervised baseline by 18.35 percentage points; with fine-tuning, F1 further improves to 93.65%, attaining a normalized edit distance of 0.93.

0 citationsRead paper

DMP-3DAD: Cross-Category 3D Anomaly Detection via Realistic Depth Map Projection with Few Normal Samples

Feb 11, 2026

This work addresses the challenge of cross-category 3D point cloud anomaly detection using only a few normal samples by proposing the first training-free, general-purpose framework. The method projects 3D point clouds into multi-view realistic depth maps and leverages a frozen CLIP vision encoder to extract features, enabling anomaly identification through weighted feature similarity without any category-specific fine-tuning or adaptation. Experimental results demonstrate that the approach achieves state-of-the-art performance under few-shot settings on the ShapeNetPart dataset, significantly enhancing the generality, practicality, and robustness of cross-category 3D anomaly detection.

0 citationsRead paper
Recent publications

Latest Papers

Visual Information Extraction from Documents via Classification-Guided Large Vision-Language Models

Jul 21, 2026

This work addresses the challenge of information extraction from visually rich documents, where diverse layouts and real-world interferences—such as stamps and low contrast—hinder performance. Existing approaches typically rely on extensive annotated data and layout-specific training. To overcome these limitations, the authors propose a classification-guided large vision-language model framework that decouples document-type classification from content extraction and leverages in-context learning to dynamically construct prompts, enabling zero-shot information extraction across multiple document types. Built upon Qwen2.5-VL-7B, the method employs conditional computation to reduce task uncertainty. Evaluated on a real-world tendering dataset comprising 16 certificate categories, it achieves a zero-shot F1 score of 86.43%, outperforming a strong supervised baseline by 18.35 percentage points; with fine-tuning, F1 further improves to 93.65%, attaining a normalized edit distance of 0.93.

0 citationsRead paper

DMP-3DAD: Cross-Category 3D Anomaly Detection via Realistic Depth Map Projection with Few Normal Samples

Feb 11, 2026

This work addresses the challenge of cross-category 3D point cloud anomaly detection using only a few normal samples by proposing the first training-free, general-purpose framework. The method projects 3D point clouds into multi-view realistic depth maps and leverages a frozen CLIP vision encoder to extract features, enabling anomaly identification through weighted feature similarity without any category-specific fine-tuning or adaptation. Experimental results demonstrate that the approach achieves state-of-the-art performance under few-shot settings on the ShapeNetPart dataset, significantly enhancing the generality, practicality, and robustness of cross-category 3D anomaly detection.

0 citationsRead paper