Institution profile

Zhongke Fanyu Technology Co., Ltd

Industry researchasia · cn
Research library8linked papers
Opportunities0open roles
Selected work

Representative Papers

ICDAR 2025 Competition on End-to-End Document Image Machine Translation Towards Complex Layouts

Mar 10, 2026

This work addresses the challenging task of end-to-end cross-lingual translation of document images with complex layouts by establishing the first systematic benchmark that jointly models textual semantics and page layout. The study introduces a dual-track framework comprising OCR-free and OCR-based approaches, accommodating both small and large language models, and integrates (optionally) optical character recognition with multimodal natural language processing to formulate a unified paradigm for document image translation. The benchmark attracted participation from 69 teams, yielding 27 valid submissions, and empirical results demonstrate that large models exhibit significant advantages and strong potential for this task.

0 citationsRead paper

PromptDLA: A Domain-aware Prompt Document Layout Analysis Framework with Descriptive Knowledge as a Cue

Mar 10, 2026

This work addresses the limited generalization of existing document layout analysis methods in cross-domain settings, which stems from their neglect of inherent structural differences among datasets. To overcome this, we propose a domain-aware prompting mechanism that leverages descriptive knowledge as domain-specific prior cues to dynamically generate tailored prompts, guiding the model to focus on salient layout features. By integrating prompt learning with a layout analysis architecture, our approach enables end-to-end domain-adaptive training. Extensive experiments on multiple benchmark datasets—including DocLayNet, PubLayNet, M6Doc, and D⁴LA—demonstrate that our method significantly outperforms current state-of-the-art approaches, achieving superior performance across diverse domains.

0 citationsRead paper

Improving MLLM's Document Image Machine Translation via Synchronously Self-reviewing Its OCR Proficiency

Jul 11, 2025

This work addresses catastrophic forgetting of monolingual capabilities—particularly OCR—induced by supervised fine-tuning in multimodal large language models (MLLMs) for document image machine translation (DIMT). We propose Synchronous Self-Check Tuning (SSCT), a novel fine-tuning paradigm that explicitly incorporates OCR text generation as an intermediate step within the translation pipeline. SSCT leverages the model’s own OCR output as self-supervised signal to jointly optimize cross-modal understanding and cross-lingual translation. Its core innovation lies in emulating “bilingual cognitive advantage” via a structured prompting mechanism that synchronously activates and preserves OCR capability during translation training. Experiments on multiple DIMT benchmarks demonstrate that SSCT significantly improves translation quality (BLEU +3.2) while maintaining or even enhancing OCR accuracy (CER −1.8%), effectively mitigating multi-task interference and enabling synergistic generalization across modalities and tasks.

0 citationsRead paper

Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation

Jul 10, 2025

Document Image Machine Translation (DIMT) suffers from insufficient training data and complex image-text modality coupling, leading to poor generalization. To address this, we propose M4Doc—a novel framework that establishes an alignment mechanism between a unimodal image encoder and a multimodal large language model (MLLM). During pretraining, M4Doc leverages the MLLM to inject joint image-text knowledge, enabling cross-modal fusion of visual and textual representations. At inference, only a lightweight image encoder is required—eliminating the need for MLLM invocation—thus ensuring efficiency and practical deployability. M4Doc is trained end-to-end on large-scale document image data. Experiments demonstrate substantial improvements in translation quality across multiple benchmarks, with particularly strong generalization to out-of-domain scenarios and documents featuring complex layouts.

0 citationsRead paper
Recent publications

Latest Papers

ICDAR 2025 Competition on End-to-End Document Image Machine Translation Towards Complex Layouts

Mar 10, 2026

This work addresses the challenging task of end-to-end cross-lingual translation of document images with complex layouts by establishing the first systematic benchmark that jointly models textual semantics and page layout. The study introduces a dual-track framework comprising OCR-free and OCR-based approaches, accommodating both small and large language models, and integrates (optionally) optical character recognition with multimodal natural language processing to formulate a unified paradigm for document image translation. The benchmark attracted participation from 69 teams, yielding 27 valid submissions, and empirical results demonstrate that large models exhibit significant advantages and strong potential for this task.

0 citationsRead paper

PromptDLA: A Domain-aware Prompt Document Layout Analysis Framework with Descriptive Knowledge as a Cue

Mar 10, 2026

This work addresses the limited generalization of existing document layout analysis methods in cross-domain settings, which stems from their neglect of inherent structural differences among datasets. To overcome this, we propose a domain-aware prompting mechanism that leverages descriptive knowledge as domain-specific prior cues to dynamically generate tailored prompts, guiding the model to focus on salient layout features. By integrating prompt learning with a layout analysis architecture, our approach enables end-to-end domain-adaptive training. Extensive experiments on multiple benchmark datasets—including DocLayNet, PubLayNet, M6Doc, and D⁴LA—demonstrate that our method significantly outperforms current state-of-the-art approaches, achieving superior performance across diverse domains.

0 citationsRead paper

Improving MLLM's Document Image Machine Translation via Synchronously Self-reviewing Its OCR Proficiency

Jul 11, 2025

This work addresses catastrophic forgetting of monolingual capabilities—particularly OCR—induced by supervised fine-tuning in multimodal large language models (MLLMs) for document image machine translation (DIMT). We propose Synchronous Self-Check Tuning (SSCT), a novel fine-tuning paradigm that explicitly incorporates OCR text generation as an intermediate step within the translation pipeline. SSCT leverages the model’s own OCR output as self-supervised signal to jointly optimize cross-modal understanding and cross-lingual translation. Its core innovation lies in emulating “bilingual cognitive advantage” via a structured prompting mechanism that synchronously activates and preserves OCR capability during translation training. Experiments on multiple DIMT benchmarks demonstrate that SSCT significantly improves translation quality (BLEU +3.2) while maintaining or even enhancing OCR accuracy (CER −1.8%), effectively mitigating multi-task interference and enabling synergistic generalization across modalities and tasks.

0 citationsRead paper

Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation

Jul 10, 2025

Document Image Machine Translation (DIMT) suffers from insufficient training data and complex image-text modality coupling, leading to poor generalization. To address this, we propose M4Doc—a novel framework that establishes an alignment mechanism between a unimodal image encoder and a multimodal large language model (MLLM). During pretraining, M4Doc leverages the MLLM to inject joint image-text knowledge, enabling cross-modal fusion of visual and textual representations. At inference, only a lightweight image encoder is required—eliminating the need for MLLM invocation—thus ensuring efficiency and practical deployability. M4Doc is trained end-to-end on large-scale document image data. Experiments demonstrate substantial improvements in translation quality across multiple benchmarks, with particularly strong generalization to out-of-domain scenarios and documents featuring complex layouts.

0 citationsRead paper