FS-DAG: Few Shot Domain Adapting Graph Networks for Visually Rich Document Understanding
Addressing key challenges in few-shot visual-rich document understanding (VRDU)—including poor cross-domain transferability, low robustness to OCR noise, and constrained deployment resources—this paper proposes a modular few-shot domain-adaptive graph network architecture. The method supports dynamic switching between language- and vision-based backbones, and integrates graph neural networks, multimodal feature alignment, domain-adaptive knowledge distillation, and OCR-robust encoding. With fewer than 90M parameters, it achieves both lightweight design and strong generalization. On information extraction tasks, the model significantly accelerates convergence and improves accuracy, consistently outperforming existing state-of-the-art approaches. It provides a scalable, highly robust solution to real-world issues such as OCR errors, spelling variants, and domain shift—enabling effective adaptation under limited supervision and noisy inputs.