DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student Distillation

📅 2026-08-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the inefficiency of current visual document retrieval (VDR) systems, which rely on billion-parameter models leading to slow indexing and high costs, and lack effective end-to-end compact single-vector solutions. The authors propose DistilVDR-HiRes, the first asymmetric dual-student distillation framework that trains a lightweight 524M-parameter system from an 8B-parameter teacher model without requiring relevance labels or negative sampling. Distillation is achieved solely through a frozen teacher-guided pairwise cosine alignment loss. The approach retains a lightweight 70M-parameter query encoder while enhancing visual document representations, integrating tile budget control and single-vector indexing. Evaluated on ViDoRe v1+v2+v3, it achieves an average NDCG@5 of 61.74—86.9% of the teacher’s performance—while reducing index size by 15.6× and accelerating index construction by an order of magnitude, substantially outperforming all reproduced sub-billion-parameter baselines.
📝 Abstract
Visual document retrieval (VDR) is dominated by multi-billion-parameter models that are slow to index at full corpus scale and expensive to serve. Prior compression routes either train a smaller multi-vector encoder from scratch or distil only the query side; neither yields a compact single-vector retriever end-to-end. We present DistilVDR, a 524M end-to-end VDR system distilled bilaterally from a single 8B vision-language teacher under a pointwise cosine alignment loss. All supervision comes from the frozen teacher's embedding space, which was itself trained with relevance supervision, so the student objective needs no relevance labels, negative sampling, or contrastive term. We match VDR's text-query and image-document input asymmetry with an asymmetric encoder-only student that concentrates visual capacity on the document side and keeps the query side at 70M parameters. We release two variants that share the same encoders and training and differ only in the document encoder's visual-tile budget: DistilVDR-HiRes attains 61.74 average NDCG@5 on ViDoRe v1+v2+v3 (86.9% of the 8B teacher) and leads every reproduced sub-1B baseline on the high-resolution-sensitive v3 benchmark, while DistilVDR-Fast attains 59.98 at a 3 times smaller visual-token budget. Both variants store one million documents in a 15.6 times smaller index than the strongest sub-1B multi-vector baseline and index the corpus an order of magnitude faster. The code is available at https://github.com/Ryenhails/NanoVDR.
Problem

Research questions and friction points this paper is trying to address.

Visual Document Retrieval
Model Compression
End-to-End Retrieval
Large Vision-Language Models
Efficient Indexing
Innovation

Methods, ideas, or system contributions that make the work stand out.

dual-student distillation
asymmetric encoder
single-vector retriever
visual document retrieval
knowledge distillation
🔎 Similar Papers
Z
Zhuchenyang Liu
Aalto University, Finland
Z
Ziyi Wang
Independent Researcher, Netherlands
Y
Yao Zhang
Aalto University, Finland
Yu Xiao
Yu Xiao
Associate Professor, Aalto University, Finland
extended realitywearable sensing/hapticsvideo analyticedge/fog computingsmart manufacturing