ViDoRe V3: A Comprehensive Evaluation of Retrieval Augmented Generation in Complex Real-World Scenarios
Existing retrieval-augmented generation (RAG) evaluation benchmarks struggle to address real-world challenges such as multi-document synthesis, visual understanding, and fine-grained source attribution. To bridge this gap, this work introduces the first comprehensive multimodal RAG benchmark that integrates visual content, cross-document reasoning, and multilingual support. The benchmark comprises 26,000 visually rich document pages spanning ten specialized domains and 3,099 human-verified queries, accompanied by high-quality annotations for retrieval relevance, bounding box localization, and reference answers. Systematic evaluation reveals that vision-aware retrievers significantly outperform purely text-based approaches, and late interaction with re-ranking further enhances performance. Nevertheless, current models still exhibit notable deficiencies in interpreting non-textual elements, answering open-ended questions, and achieving fine-grained visual grounding.